screen-reader-testing
Test web applications with screen readers including VoiceOver, NVDA, and JAWS. Use when validating screen reader compatibility, debugging accessibility issues,…
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
$ npx -y skills add wshobson/agents --skill eval-harness-first --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/eval-harness-firstContext preview
The summary Claude sees to decide when to auto-load this skill.
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
name: eval-harness-first description: Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
The Phase 0 gate for the whole plugin: `finetuning-method-selection` and every downstream skill assume this harness exists before a training config gets written. The harness is not a run-end side artifact — it is the data-curation engine. The same labeled traces that build the goldens feed training data, minus an explicit holdout.
**Input:** production/agent traces if they exist, or a task spec if they don't, plus labelers willing to grade ≥100 examples. **Output format:** the `eval/` directory below — goldens, graders, drift suite, and the base-model baseline that later phases gate on.
No eval harness, no fine-tune. Skip to a training config and there is nothing to measure against, nothing to catch regressions, and no labeled data to train on. The flywheel:
1. **Collect traces** — production/agent spans, or synthetic tasks if none exist yet. 2. **Error analysis** — open coding on ≥100 traces, axial coding into 4–8 failure buckets. 3. **One grader per bucket** — deterministic first; calibrated LLM-judge only for genuinely subjective criteria. 4. **Prioritize** by frequency × severity × value. 5. **The labeled traces feed dataset curation, minus an explicit holdout.** Every `eval/goldens.jsonl` ID stays excluded from training data by ID. 6. **Train.** 7. **Re-run the same harness** on the checkpoint — not a different, looser one. 8. **Drift detection feeds back to step 2** — new production failure modes re-open error analysis.
Steps 2–4 build the harness; steps 5–8 are why it must exist first — it is both the training data source and the checkpoint's exit gate.
analysis — open coding on ≥100 real traces (read them, tag failures in your own words, no fixed taxonomy yet), then axial coding to collapse those tags into 4–8 named failure buckets. Fewer than 4 means the coding pass was too shallow; more than 8 means buckets need merging. **Exception:** single-failure-surface tasks (e.g. strict-schema extraction) may land at 1–2 buckets with per-field sub-metrics inside one grader — don't invent artificial splits with no evidence behind them.
dimension-based generation — enumerate the axes that matter (task type, difficulty, edge case, persona) and sample the cross-product; free- generated prompts cluster around whatever's easiest to write.
`eval/goldens.jsonl`, diff it in review, tag it per release. It doubles as the CI regression suite.
One grader per failure bucket from error analysis — not one for the whole eval set. A single blended score hides which bucket regressed.
or execution checks are cheaper, reproducible, and need no calibration.
criteria** — tone, faithfulness, "which response is better" — where no deterministic check can express it.
scale is noisier to calibrate and harder to apply consistently; collapse to pass/fail.
over generate-and-extract** — a tight token budget makes generate-and-extract parse-brittle for models that preamble, conflating format compliance with the knowledge being measured. Templates for all four grader shapes and this scoring note: `references/grader-templates.md`.
Any bucket routed to an LLM-judge needs calibration before its verdicts count for anything beyond exploration — a hard prerequisite, not a nice-to-have. **N/A when no bucket routes to a judge** — an all-deterministic harness has nothing to calibrate; state that rather than leaving this section unaddressed.
test** (report once, no re-touching after).
number — a judge can hit 90% by always saying "pass" on a skewed set.
recalibrate on judge-model change, quarterly regardless.
than the model under test.**
**advisory-only** — flags for human review, never gates a promotion. Full protocol, bias correction, and recalibration checklist: `references/judge-calibration.md`.
Before Phase 1 (method selection) starts, run the full harness — goldens plus the capability-drift suite — against the unmodified base model. This is the number every later checkpoint gets compared against.
`eval/baseline-<model>.json` is the gate token. No baseline file, no comparison basis for `checkpoint-promotion` — a checkpoint that "looks better" against nothing measured isn't a finding.
eval/
├── goldens.jsonl # labeled traces + synthetic goldens, versioned
├── graders/ # one module per failure bucket
│ ├── schema_compliance.py
│ ├── exact_match.py
│ └── rubric_judge.py
├── drift-suite.yaml # frozen benchmarks + 200-500 domain-adjacent items
└── baseline-<model>.json # gate token: harness + drift suite vs the base model
runs/
└── <run-id>/
└── results.json # per-run harness output, one per checkpoint`eval/` persists across runs and lives outside `runs/` — the fixed measuring stick, not a run artifact. `runs/` is disposable; `eval/` is not. Never let a run script write into `eval/`. **Canonical location:**
Production-ready agentic workflow building blocks: 94 plugins, 202 agents, 183 skills, 105 commands — built for Claude Code and consumed natively by OpenAI Codex CLI, Cursor, OpenCode, the Antigravity CLI, GitHub Copilot, and Pi from a single Markdown source.
Repo: wshobson/agents
Test web applications with screen readers including VoiceOver, NVDA, and JAWS. Use when validating screen reader compatibility, debugging accessibility issues,…
Conduct WCAG 2.2 accessibility audits with automated testing, manual verification, and remediation guidance. Use when auditing websites for accessibility,…
Coordinate parallel code reviews across multiple quality dimensions with finding deduplication, severity calibration, and consolidated reporting. Use this…
Debug complex issues using competing hypotheses with parallel investigation, evidence collection, and root cause arbitration. Use this skill when debugging…
Coordinate parallel feature development with file ownership strategies, conflict avoidance rules, and integration patterns for multi-agent implementation. Use…
Decompose complex tasks, design dependency graphs, and coordinate multi-agent work with proper task descriptions and workload balancing. Use this skill when…