model-onboarding
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says:
$ npx -y skills add rlaope/oh-my-hermes --skill omh-inference-serving --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/omh-inference-servingContext preview
The summary Claude sees to decide when to auto-load this skill.
[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says:
name: "omh-inference-serving"
description: "[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says: inference-serving, inference serving, serve this model, serve the model, model serving, serving endpoint, vllm, llama.cpp."
metadata:
hermes:
tags: [workflow, oh-my-hermes, operations]
category: operations
phase: inference-serving
role: operator
quality_tier: observed-command-gatedThis is a Hermes-native `inference-serving` workflow skill.
`inference-serving` exists so serving an LLM runs as one decided, gated, measured process instead of scattered flag folklore: the engine choice is a table, the deployment is an idempotent runbook whose only completion evidence is the observed verification, and the benchmark speaks the standard metric vocabulary.
Good example:
Bad example:
Use when a model needs to be served - engine and quantization chosen, docker or Kubernetes deployment prepared as a gated runbook, or the endpoint measured with the TTFT/TPOT/ITL/goodput protocol - and the user wants the process, not an ad-hoc command guess.
Strong routing signals: `inference-serving`, `inference serving`, `serve this model`, `serve the model`, `model serving`, `serving endpoint`, `vllm`, `llama.cpp`, `llama cpp`, `serve with vllm`, `deploy vllm`, `vllm deployment`, `serving benchmark`, `benchmark the endpoint`, `prefix caching benchmark`, `gguf quantization`, `which quantization`, `모델 서빙`, `모델 서빙해줘`, `모델 배포해서 서빙`, `서빙 벤치마크`, `vllm 배포`, `vllm 서빙`, `추론 서버 띄워줘`, `모델 띄워줘`
Category: `operations` Phase: `inference-serving` Hermes role: `operator` Quality tier: `observed-command-gated` Reasoning demand: `light`
Quality bar:
Handoff policy:
Keep engine/quantization decisions, runbook preparation, and benchmark design in Hermes; the commands run through the operator's terminal with observed evidence, and repository changes (deploy manifests, benchmark harnesses) are coding work for the selected executor lane. A runbook or benchmark plan is prepared_not_observed until its commands' results are seen.
Required inputs:
English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.
Repo: rlaope/oh-my-hermes
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
Review oh-my-hermes pull requests that have not been reviewed at their current head commit.…
Backfill labels across oh-my-hermes issues and pull requests. Run manually to sweep…
[omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus,…
[omh] Hermes badges unlocked and achievement progress: achievements observation: summarize…
[omh] Technical proposal facing adversarial scrutiny: independent perspectives attack a…