architecture
Use when the user asks to improve architecture, find refactoring opportunities, surface…
Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete,
$ npx -y skills add Kanevry/session-orchestrator --skill test-audit --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/test-auditContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete,
name: test-audit description: > Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete, merge, or repair carries written evidence and a caught mutation. Triggered by "audit the tests", "prune the test suite", "remove low-value tests", "test diet", "is this test worth adding", /test-audit. user-invocable: true argument-hint: "[gate|audit|campaign] [scope] [--goal \"<goal>\"]" model: inherit
A test earns its place by catching a regression nothing else catches. This skill applies that bar at three depths. It complements `.claude/rules/test-value.md` (TV-001..TV-005) and the repo's test-quality / test-hygiene rules; it replaces none of them.
Rule references below (`test-value.md`, `receiving-review.md`, `parallel-sessions.md`, `bash-harness-pitfalls.md`) name the repo's `.claude/rules/`. A consumer repo often lacks them; then read them from the plugin's `rules/always-on/`.
Invoked as `/test-audit [gate|audit|campaign] [scope] [--goal "<goal>"]` with arguments: **$ARGUMENTS**.
State the goal in the first output and repeat it in the report. With no `--goal`, use:
> Remove at least 20 % of the least useful tests, measured in test lines. Total coverage stays within 2 percentage points of M0, and no production file drops below its own M0 coverage.
Test lines count only for files the CI actually runs; tracked test files outside that scope are reported on their own line, as dead or as live outside the include, and never counted toward the goal ([references/campaign.md](references/campaign.md) § Measurements).
The goal exists because a bare "clean up" stops far too early. It never licenses a deletion without evidence. What counts toward it, and what is reported beside it (same-file reshaping, mock ballast, F/N growth, the O sum), is fixed in [references/campaign.md](references/campaign.md) § Goal accounting. A goal that is only reachable by breaking the evidence rules below is not reached: the report says which rule blocked it and where the audit stopped. An audit that ends with zero changes is valid when the ledger shows why.
Before adding any test, answer all four. A missing answer means the test is not written yet.
1. Which observable behaviour, invariant, or independent contract does it protect? 2. Which credible regression turns it red? 3. Why does no existing test catch that regression? Grep first (TV-004); extend a table case or shared fixture instead of adding a near-duplicate. 4. Does it need a production seam (export, flag, hook, injection parameter) that no production caller uses? Then no: move the test to the real boundary. An export or parameter with one production caller is API when the package's public entry (`exports`, index, another repo) reaches it, and a seam otherwise.
Then check it against the junk patterns below; a match fails the gate unless the keep list names the contract it alone guards. A regression test must be shown red on the pre-fix code, for the intended reason, before the fix lands; one that never failed proves the mock, not the fix. One regression at the owning boundary covers the bug; do not replay it at every layer. `no-tests-needed: <reason>` is a success outcome (TV-001).
Judge each test by what its assertions can catch, never by its name.
**Junk patterns** (candidates for D, C, or F):
**Keep list** (the counterweight to TV-002): a test that is the independent proof for a public API, protocol, a config value production reads (a config field whose only reader is the test is a seam, not a contract), migration, storage format, security control, release step, or data-protection rule stays. So do observable call ordering and regressions with a credible failure mode. Static or slow is n
Give your agents a working rhythm. Your coding agents get an agreed plan, a check after each round of changes and a handover to the next session.
Repo: Kanevry/session-orchestrator
Use when the user asks to improve architecture, find refactoring opportunities, surface…
Use this skill when running an autonomous session-orchestration loop. Chains session-start →…
Use this skill when scaffolding the minimum repository structure required by…
Use when you have a feature idea but the scope or UX is still ambiguous — runs a lightweight…
Use when detecting drift between CLAUDE.md (or AGENTS.md, the Codex CLI alias) / _meta…