screen-reader-testing
Test web applications with screen readers including VoiceOver, NVDA, and JAWS. Use when validating screen reader compatibility, debugging accessibility issues,…
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
$ npx -y skills add wshobson/agents --skill checkpoint-promotion --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/checkpoint-promotionContext preview
The summary Claude sees to decide when to auto-load this skill.
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
name: checkpoint-promotion description: Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
The Phase 5 gate for the whole plugin: a checkpoint that trains cleanly and beats its task metric still doesn't ship without clearing all four stages below. `eval-harness-first` built the suite re-run here — this skill is where that suite's baseline decides something.
**Input:** a trained checkpoint, `eval/baseline-<model>.json` from `eval-harness-first`, and the frozen `eval/drift-suite.yaml`. **Output format:** `promotion-report.md` — the four-stage evidence plus a terminal `PROMOTE` or `REJECT` verdict that `/finetune` Phase 5 and `/promote-checkpoint` consume directly.
Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a **deterministic** arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.
1. **Data-quality gate.** Before any eval touches the checkpoint: dedup the training set, check for eval-goldens leakage (the exact failure `trace-to-training-data`'s Hygiene section exists to prevent), and scan for label noise. A checkpoint trained on leaked goldens invalidates every later stage. 2. **Held-out + frozen capability-drift suite.** Re-run `eval-harness-first`'s `eval/drift-suite.yaml` — MMLU/GSM8K/IFEval plus 200–500 domain-adjacent items — against the checkpoint and diff against `baseline-<model>.json` per benchmark against the Drift Budget table below. 3. **Paired arena vs. base.** Position-randomized judge, checkpoint vs. base model, same prompts — or the deterministic paired-comparison variant in `references/gate-templates.md` when every grader in the harness is deterministic (no LLM-judge; position randomization N/A there). **A holdout win that loses the live arena does not ship** — stage-2 numbers and stage-3 judgments must agree; a win on frozen goldens and a loss in paired comparison is a real signal, not a discrepancy to explain away. 4. **Canary.** 5–10% stratified rollout with auto-rollback for any checkpoint reaching production traffic. **Local-only users stop at stage 3** — skipping stage 4 for a local deployment is the correct stopping point, not a shortcut.
| Drift (pts) | Verdict | |---|---| | ≤1 | Noise — proceed | | 2–5 | Rerun with seed variation before deciding | | >5 | **HARD FAIL** — no exception for task gains |
The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.
**Item count derives from the budget, not convenience:** the strict n for a half-width under half the 5pt hard-fail threshold is ~1,300 at typical accuracy (p≈0.7); n=200 is a pragmatic floor (±6pt half-width at that same p, n=50 ±13pt) — report the half-width with every verdict, and treat a margin smaller than it as `REJECT (uncertain)`, not PASS/HARD FAIL. Full math and a 5-run cautionary example: `references/gate-templates.md`.
**RERUN is not a verdict.** A 2–5pt drift only ever produces a `PROMOTE` or `REJECT` after the seed-variation rerun completes — `PROMOTE` requires landing back at ≤1pt (noise); any rerun still >1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard `REJECT`. No report may reach the Verdict section with stage 2 still showing `RERUN`.
Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:
unmanaged** — no replay, no regularization.
— some replay or a conservative LR.
disciplined case.
mix is the standard mitigation** — blend general- domain data into training rather than target-task data alone.
If a checkpoint hits the >5pt hard fail in stage 2, work this escalation ladder in order — the one canonical order this skill and `references/gate-templates.md` both point to:
1. **Adjust the replay-mix fraction — swap rows, don't add them** (adding confounds fraction with total optimizer steps). Dose is not monotonic at small-run scale (<~100 steps) — re-check drift after any swap. 2. **Lower the learning rate.** 3. **Fewer epochs.** 4. **A smaller LoRA rank** — the same rank/LR levers `lora-qlora-recipes` and `preference-optimization` tune for the training run, applied here in reverse.
This order is a default, not a law: **remediation guidance from a single before/after run pair is a hypothesis** — label it low-confidence once any lever produces a reversal, and prefer a seed-variation repeat over trusting the next rung blindly. A lever that clears the drift breach but drops a success-criterion metric below target is a two-sided tradeoff for a human, not a reason to keep descending the ladder. Full reasoning and the 5-run trajectory behind both caveats: `references/gate-templates.md`.
**Disclose drift-suite instruction reuse.** A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.
`promotion-report.md` covers all four stages as sections and **must end with a terminal verdict: `PROMOTE` or `REJECT`**, the eviden
Production-ready agentic workflow building blocks: 94 plugins, 202 agents, 183 skills, 105 commands — built for Claude Code and consumed natively by OpenAI Codex CLI, Cursor, OpenCode, the Antigravity CLI, GitHub Copilot, and Pi from a single Markdown source.
Repo: wshobson/agents
Test web applications with screen readers including VoiceOver, NVDA, and JAWS. Use when validating screen reader compatibility, debugging accessibility issues,…
Conduct WCAG 2.2 accessibility audits with automated testing, manual verification, and remediation guidance. Use when auditing websites for accessibility,…
Coordinate parallel code reviews across multiple quality dimensions with finding deduplication, severity calibration, and consolidated reporting. Use this…
Debug complex issues using competing hypotheses with parallel investigation, evidence collection, and root cause arbitration. Use this skill when debugging…
Coordinate parallel feature development with file ownership strategies, conflict avoidance rules, and integration patterns for multi-agent implementation. Use…
Decompose complex tasks, design dependency graphs, and coordinate multi-agent work with proper task descriptions and workload balancing. Use this skill when…