model-onboarding
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor
$ npx -y skills add rlaope/oh-my-hermes --skill omh-agent-evaluation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/omh-agent-evaluationContext preview
The summary Claude sees to decide when to auto-load this skill.
[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor
name: "omh-agent-evaluation"
description: "[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude."
metadata:
hermes:
tags: [workflow, oh-my-hermes, operations]
category: operations
phase: agent-evaluation
role: operator
quality_tier: agent-eval-gatedThis is a Hermes-native `agent-evaluation` workflow skill.
`agent-evaluation` gives OMH a way to improve executor choice empirically, not by vibes, while preserving executor-neutral product language across Codex, Claude Code, Hermes, and generic runtimes.
Good example:
Bad example:
Use when Hermes should design or summarize a fair comparison of Codex, Claude Code, Hermes coding, or generic executors for a bounded task set.
Strong routing signals: `agent-evaluation`, `agent evaluation`, `agent eval`, `agent benchmark`, `executor evaluation`, `executor benchmark`, `compare agents`, `compare codex claude`, `agent tournament`, `which agent is better`, `에이전트 평가`, `에이전트 비교`, `실행자 평가`, `코덱스 클로드 비교`
Category: `operations` Phase: `agent-evaluation` Hermes role: `operator` Quality tier: `agent-eval-gated` Reasoning demand: `light`
Quality bar:
Handoff policy:
Keep evaluation design and scoring in Hermes. Actual executor runs, costs, timings, tool calls, code edits, and review results must come from observed runtime or supplied artifacts.
Required inputs:
Expected outputs:
Artifact expectations:
Safety ru
English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.
Repo: rlaope/oh-my-hermes
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
Review oh-my-hermes pull requests that have not been reviewed at their current head commit.…
Backfill labels across oh-my-hermes issues and pull requests. Run manually to sweep…
[omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus,…
[omh] Hermes badges unlocked and achievement progress: achievements observation: summarize…
[omh] Technical proposal facing adversarial scrutiny: independent perspectives attack a…