Skip to content
Development
Skill

/checkpoint-promotion

Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.

From plugin
wshobson-agents
40k183 skills137 agents93 commands
Install
$ npx -y skills add wshobson/agents --skill checkpoint-promotion --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/checkpoint-promotion

Context preview

The summary Claude sees to decide when to auto-load this skill.

Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.

SKILL.md

checkpoint-promotion.SKILL.md
name: checkpoint-promotion
description: Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.

Checkpoint Promotion

The Phase 5 gate for the whole plugin: a checkpoint that trains cleanly and beats its task metric still doesn't ship without clearing all four stages below. `eval-harness-first` built the suite re-run here — this skill is where that suite's baseline decides something.

**Input:** a trained checkpoint, `eval/baseline-<model>.json` from `eval-harness-first`, and the frozen `eval/drift-suite.yaml`. **Output format:** `promotion-report.md` — the four-stage evidence plus a terminal `PROMOTE` or `REJECT` verdict that `/finetune` Phase 5 and `/promote-checkpoint` consume directly.

The Four-Stage Gate

Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a **deterministic** arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.

1. **Data-quality gate.** Before any eval touches the checkpoint: dedup the training set, check for eval-goldens leakage (the exact failure `trace-to-training-data`'s Hygiene section exists to prevent), and scan for label noise. A checkpoint trained on leaked goldens invalidates every later stage. 2. **Held-out + frozen capability-drift suite.** Re-run `eval-harness-first`'s `eval/drift-suite.yaml` — MMLU/GSM8K/IFEval plus 200–500 domain-adjacent items — against the checkpoint and diff against `baseline-<model>.json` per benchmark against the Drift Budget table below. 3. **Paired arena vs. base.** Position-randomized judge, checkpoint vs. base model, same prompts — or the deterministic paired-comparison variant in `references/gate-templates.md` when every grader in the harness is deterministic (no LLM-judge; position randomization N/A there). **A holdout win that loses the live arena does not ship** — stage-2 numbers and stage-3 judgments must agree; a win on frozen goldens and a loss in paired comparison is a real signal, not a discrepancy to explain away. 4. **Canary.** 5–10% stratified rollout with auto-rollback for any checkpoint reaching production traffic. **Local-only users stop at stage 3** — skipping stage 4 for a local deployment is the correct stopping point, not a shortcut.

Drift Budget

| Drift (pts) | Verdict | |---|---| | ≤1 | Noise — proceed | | 2–5 | Rerun with seed variation before deciding | | >5 | **HARD FAIL** — no exception for task gains |

The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.

**Item count derives from the budget, not convenience:** the strict n for a half-width under half the 5pt hard-fail threshold is ~1,300 at typical accuracy (p≈0.7); n=200 is a pragmatic floor (±6pt half-width at that same p, n=50 ±13pt) — report the half-width with every verdict, and treat a margin smaller than it as `REJECT (uncertain)`, not PASS/HARD FAIL. Full math and a 5-run cautionary example: `references/gate-templates.md`.

**RERUN is not a verdict.** A 2–5pt drift only ever produces a `PROMOTE` or `REJECT` after the seed-variation rerun completes — `PROMOTE` requires landing back at ≤1pt (noise); any rerun still >1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard `REJECT`. No report may reach the Verdict section with stage 2 still showing `RERUN`.

Catastrophic Forgetting

Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:

  • **~43% knowledge loss

unmanaged** — no replay, no regularization.

  • **~10% with basic management**

— some replay or a conservative LR.

  • **~3% with replay + EWC** — the

disciplined case.

  • **10–30% general-data replay

mix is the standard mitigation** — blend general- domain data into training rather than target-task data alone.

If a checkpoint hits the >5pt hard fail in stage 2, work this escalation ladder in order — the one canonical order this skill and `references/gate-templates.md` both point to:

1. **Adjust the replay-mix fraction — swap rows, don't add them** (adding confounds fraction with total optimizer steps). Dose is not monotonic at small-run scale (<~100 steps) — re-check drift after any swap. 2. **Lower the learning rate.** 3. **Fewer epochs.** 4. **A smaller LoRA rank** — the same rank/LR levers `lora-qlora-recipes` and `preference-optimization` tune for the training run, applied here in reverse.

This order is a default, not a law: **remediation guidance from a single before/after run pair is a hypothesis** — label it low-confidence once any lever produces a reversal, and prefer a seed-variation repeat over trusting the next rung blindly. A lever that clears the drift breach but drops a success-criterion metric below target is a two-sided tradeoff for a human, not a reason to keep descending the ladder. Full reasoning and the 5-run trajectory behind both caveats: `references/gate-templates.md`.

**Disclose drift-suite instruction reuse.** A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.

The Verdict

`promotion-report.md` covers all four stages as sections and **must end with a terminal verdict: `PROMOTE` or `REJECT`**, the eviden

Read more
Ships withwshobson-agents

Production-ready agentic workflow building blocks: 94 plugins, 202 agents, 183 skills, 105 commands — built for Claude Code and consumed natively by OpenAI Codex CLI, Cursor, OpenCode, the Antigravity CLI, GitHub Copilot, and Pi from a single Markdown source.

Get the whole plugin

Other skills on wshobson-agents.