Skip to content
Development
Skill

/test-audit

Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete,

BOOST
From plugin
session-orchestrator
5351 skills14 agents2 commands10 hooks
+1
Install
$ npx -y skills add Kanevry/session-orchestrator --skill test-audit --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/test-audit

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete,

SKILL.md

test-audit.SKILL.md
name: test-audit
description: >
  Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate
  before any new test is written, a focused audit of one scope, and a campaign over a whole repo or
  subsystem that removes the least useful tests under a coverage guard. Every delete, merge, or
  repair carries written evidence and a caught mutation. Triggered by "audit the tests", "prune the
  test suite", "remove low-value tests", "test diet", "is this test worth adding", /test-audit.
user-invocable: true
argument-hint: "[gate|audit|campaign] [scope] [--goal \"<goal>\"]"
model: inherit

Test Audit

A test earns its place by catching a regression nothing else catches. This skill applies that bar at three depths. It complements `.claude/rules/test-value.md` (TV-001..TV-005) and the repo's test-quality / test-hygiene rules; it replaces none of them.

Rule references below (`test-value.md`, `receiving-review.md`, `parallel-sessions.md`, `bash-harness-pitfalls.md`) name the repo's `.claude/rules/`. A consumer repo often lacks them; then read them from the plugin's `rules/always-on/`.

Invocation

Invoked as `/test-audit [gate|audit|campaign] [scope] [--goal "<goal>"]` with arguments: **$ARGUMENTS**.

  • `gate` — run the write gate on the test about to be written. Default when no mode is given and the conversation is about to add a test.
  • `audit <scope>` — one focused pass over a path, glob, or module. Default when a scope is given without a mode.
  • `campaign <scope>` — a whole repo or subsystem, split into lanes. Only on the explicit word `campaign`; read [references/campaign.md](references/campaign.md) before starting.
  • A missing scope for `audit` or `campaign` is asked for once via `AskUserQuestion`, with the largest test directories as options.

Set an ambitious goal first (audit and campaign)

State the goal in the first output and repeat it in the report. With no `--goal`, use:

> Remove at least 20 % of the least useful tests, measured in test lines. Total coverage stays within 2 percentage points of M0, and no production file drops below its own M0 coverage.

Test lines count only for files the CI actually runs; tracked test files outside that scope are reported on their own line, as dead or as live outside the include, and never counted toward the goal ([references/campaign.md](references/campaign.md) § Measurements).

The goal exists because a bare "clean up" stops far too early. It never licenses a deletion without evidence. What counts toward it, and what is reported beside it (same-file reshaping, mock ballast, F/N growth, the O sum), is fixed in [references/campaign.md](references/campaign.md) § Goal accounting. A goal that is only reachable by breaking the evidence rules below is not reached: the report says which rule blocked it and where the audit stopped. An audit that ends with zero changes is valid when the ledger shows why.

Mode 1: write gate

Before adding any test, answer all four. A missing answer means the test is not written yet.

1. Which observable behaviour, invariant, or independent contract does it protect? 2. Which credible regression turns it red? 3. Why does no existing test catch that regression? Grep first (TV-004); extend a table case or shared fixture instead of adding a near-duplicate. 4. Does it need a production seam (export, flag, hook, injection parameter) that no production caller uses? Then no: move the test to the real boundary. An export or parameter with one production caller is API when the package's public entry (`exports`, index, another repo) reaches it, and a seam otherwise.

Then check it against the junk patterns below; a match fails the gate unless the keep list names the contract it alone guards. A regression test must be shown red on the pre-fix code, for the intended reason, before the fix lands; one that never failed proves the mock, not the fix. One regression at the owning boundary covers the bug; do not replay it at every layer. `no-tests-needed: <reason>` is a success outcome (TV-001).

The value bar

Judge each test by what its assertions can catch, never by its name.

**Junk patterns** (candidates for D, C, or F):

  • no assertion, or an assertion that cannot fail (self-comparison, identity copy);
  • copied inventories, export lists, manifests, or fixtures that restate the source;
  • a test-local copy of another repo's schema or contract, checked against test-local payloads: it proves the copy, not the contract; the keeper is the synced artefact or the other repo's test;
  • exact source, import, or string greps where a behavioural check exists;
  • the expected value is produced by the helper under test;
  • the mock implements the behaviour being asserted, or one mock stands in for different APIs;
  • a negative control that passes for an unrelated reason (a different guard rejects first, or the production path never reaches the rejection); for a security check, one that only hits the early exit (wrong length, wrong prefix) and never the comparison itself is usually F;
  • the name promises more than the assertion checks;
  • an assertion that depends on the current date (a fixture expiring next quarter turns it red on its own; F with a fixed clock, and urgent);
  • prose or structure pinned in a `.md` file (TV-002c), unless the project instruction file makes that structure a rule (then R, citing it);
  • the test is the only caller of dead production code: D only together with the removal of that code (its own `code-implementer` commit, owner decision when the code is public API), otherwise O.

**Keep list** (the counterweight to TV-002): a test that is the independent proof for a public API, protocol, a config value production reads (a config field whose only reader is the test is a seam, not a contract), migration, storage format, security control, release step, or data-protection rule stays. So do observable call ordering and regressions with a credible failure mode. Static or slow is n

Read more
Ships withsession-orchestrator

Give your agents a working rhythm. Your coding agents get an agreed plan, a check after each round of changes and a handover to the next session.

Get the whole plugin

Other skills on session-orchestrator.