Skip to content
Automation
Skill

/eval-harness

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

From plugin
babysitter
1.8k200 skills3 agents21 commands1 MCP
Install
$ npx -y skills add a5c-ai/babysitter --skill eval-harness --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/eval-harness

Context preview

The summary Claude sees to decide when to auto-load this skill.

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

SKILL.md

eval-harness.SKILL.md
name: eval-harness
description: Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.
allowed-tools: Read, Write, Edit, Bash, Grep, Glob
graph:
  domains: [domain:software-engineering]
  skillAreas: [skill-area:agentic-loops, skill-area:orchestration-loop]
  workflows: [workflow:feature-development]
  topics: [topic:developer-experience]
  roles: [role:tech-lead, role:backend-engineer]
  • Define test cases with known-correct outputs
  • Run agent against each test case
  • Score: accuracy, completeness, relevance
  • Compare against baseline performance
  • Track performance over time

2. Skill Quality Testing

  • Verify skill instructions produce expected outcomes
  • Test edge cases and boundary conditions
  • Measure consistency across multiple runs
  • Check for harmful or incorrect outputs
  • Validate against ground truth

3. Regression Suite

  • Collection of previously-passing test cases
  • Run after any agent/skill modification
  • Flag regressions with before/after comparison
  • Maintain pass rate threshold (>= 95%)

4. Process Verification

  • End-to-end process execution with known inputs
  • Verify each phase produces expected outputs
  • Check task ordering and dependency satisfaction
  • Measure total execution time

Quality Scoring

Accuracy Score (0-100)

  • Correctness of output vs expected
  • Partial credit for partially correct outputs
  • Penalty for hallucinated or fabricated content

Completeness Score (0-100)

  • Coverage of required output elements
  • Missing sections flagged and scored
  • Bonus for useful additional context

Consistency Score (0-100)

  • Run same input 3 times
  • Compare outputs for semantic similarity
  • Flag inconsistencies

Composite Score

  • (accuracy * 0.4 + completeness * 0.3 + consistency * 0.3)
  • Threshold: 80 to pass

When to Use

  • After creating new agents or skills
  • After modifying existing agents or skills
  • Periodic quality audits
  • Before promoting skills to production

Agents Used

  • Used by process-level evaluation orchestrators
  • No specific agent dependency (evaluates other agents)
Read more
Ships withbabysitter

Enforce obedience on agentic workforces. Manage extremely complex workflows through deterministic, hallucination-free self-orchestration.

Get the whole plugin

Other skills on babysitter.