Skip to content

/eval-harness

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

From plugin
1.6k200 skills2 agents1 MCP
shell
$ npx -y skills add a5c-ai/babysitter --skill eval-harness --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/eval-harness
How auto-invocation works

Context preview

The summary Claude sees to decide when to auto-load this skill.

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.
Ships withbabysitter

Enforce obedience on agentic workforces. Manage extremely complex workflows through deterministic, hallucination-free self-orchestration.

Get the whole plugin, auto-invoked

Other skills on babysitter.