Skip to content
Business
Skill

/skill-audit-loop

Measure whether an existing Claude Code skill actually works, then fix it from evidence instead of by inspection. Use this whenever someone asks if a skill is working, wants to test, audit, benchmark, or evaluate a skill, says a skill is being ignored or fires when it shouldn't,

BOOST
From plugin
solo-founder-skills
25262 skills1 command
Install
$ npx -y skills add whawkinsiv/solo-founder-skills --skill skill-audit-loop --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/skill-audit-loop

Context preview

The summary Claude sees to decide when to auto-load this skill.

Measure whether an existing Claude Code skill actually works, then fix it from evidence instead of by inspection. Use this whenever someone asks if a skill is working, wants to test, audit, benchmark, or evaluate a skill, says a skill is being ignored or fires when it shouldn't,

SKILL.md

skill-audit-loop.SKILL.md
name: skill-audit-loop
description: Measure whether an existing Claude Code skill actually works, then fix it from evidence instead of by inspection. Use this whenever someone asks if a skill is working, wants to test, audit, benchmark, or evaluate a skill, says a skill is being ignored or fires when it shouldn't, wants an A/B comparison of results with and without a skill, wants to turn a real session where a skill misbehaved into a regression case, or wants to revise a SKILL.md and doesn't want to guess at the edit. Also use when a skill "seems fine" but nobody has actually checked. For drafting a brand-new skill from nothing, use skill-creator instead, then come back here once a draft exists.

Skill audit loop

Skills are prompts, and prompts fail quietly. A skill can look excellent and still be ignored half the time, or be followed exactly and still produce worse output than no skill at all. You cannot see either of those by reading the SKILL.md. You see them by running the skill against fixed tasks, comparing against a run without it, and reading what the agent actually did.

This skill runs that loop:

1. **Collect cases** — real sessions where the skill was used, plus a few constructed ones. 2. **Write checks** — before touching the skill. The checks are the specification. 3. **Run the matrix** — every case, every arm, several repetitions. 4. **Read the evidence** — pass rates, cost, and the trace of what the agent did. 5. **Revise once, verify, keep or revert** — one hypothesis per round, with the user's sign-off.

The loop is not optional at step 5. A revision that isn't re-measured is just another guess with more words in it.

Before running anything

Settle the auth question first, because it decides both what the runs cost and whether the comparison means anything.

**Pick one of two configurations.** They trade the same thing in opposite directions.

*API key, bare mode* — `"auth": "api_key"`, `"bare": true`. `--bare` skips auto-discovery of hooks, skills, plugins, MCP servers, memory, and CLAUDE.md, so the baseline arm is genuinely skill-free and the with-skill arm gets exactly one skill: the one under test, staged into an `--add-dir` directory. This is the cleanest isolation available. It bills per token, and bare mode never reads OAuth credentials, the system keychain, or `CLAUDE_CODE_OAUTH_TOKEN` — so a subscription cannot be used here at all. Bare sessions also carry a smaller tool set.

*Subscription, dedicated profile* — `"auth": "subscription"`, `"bare": false`, `"config_dir": "~/.claude-audit"`. `CLAUDE_CONFIG_DIR` gives Claude Code a separate profile with its own credentials, settings, skills, hooks, and MCP servers. Log into it once and every run is isolated *and* covered by the subscription:

CLAUDE_CONFIG_DIR=~/.claude-audit claude auth login

Two things to watch. `CLAUDE_CONFIG_DIR` is thinly documented and there are open reports of it still creating local `.claude/` directories in the workspace, so confirm the isolation with a smoke run rather than assuming it. And if `ANTHROPIC_API_KEY` is set anywhere in the environment it takes precedence over a subscription — you can believe you're spending rate limits and be spending dollars. The harness strips it from the child environment when `auth` is `subscription`, but check `unset ANTHROPIC_API_KEY` in your own shell too.

**Either way, it costs something.** A modest suite (5 cases x 2 arms x 3 reps) is 30 agent sessions. On an API key that's a bill; the harness records `total_cost_usd` per run and reports the total. On a subscription it's rate limits, drawn from the same pool as your own work, so keep `--jobs` low or you'll spend the afternoon throttled. If the budget is tight, cut cases before cutting repetitions — see below, because 1 rep is close to worthless.

**Never run without isolation.** If bare mode is off and no `config_dir` is set, the baseline arm loads every skill you have installed, and the comparison measures nothing. The harness warns about this; don't wave it through. At absolute minimum, keep the skill under test out of `~/.claude/skills` and manage it from a repo directory, so the only difference between arms is the `--add-dir`.

**Smoke test before spending.** `scripts/run_matrix.py --smoke` runs one case, one arm, one rep, and prints the tool list from the session's init event. That catches a broken auth setup and a tool the skill needs but the session doesn't have — bare sessions get Bash, file read, and file edit, so a skill depending on web fetch or MCP will fail there for reasons that have nothing to do with the skill. Say which configuration you used in the report, because it changes how much the numbers mean.

Running the scripts

Every `scripts/...` path below is relative to the folder that holds this SKILL.md, not to your project. Resolve it against that folder before you run it. If your agent reports `No such file or directory`, it used the wrong working directory: prefix the path with the folder this file was loaded from.

Stage 1 — Collect cases

Cases from real failures beat cases you invented. An invented case tests whether the agent can solve a puzzle; a real one tests the thing that actually went wrong.

Start with the field log:

python3 scripts/scan_field.py --skill <skill-name> --days 30 --out cases/field.md

This scans local Claude Code session transcripts for sessions where the skill was loaded and pulls out every user turn that came after it. Read them with the user. The turns where they had to correct the agent — "no, put it in the other format", "you skipped the part where…" — are the highest-value cases available, because each one is a recorded instance of the skill failing to carry an instruction, with the correct answer supplied by the user for free.

The script ranks candidates but does not judge them. Ask the user which of those exchanges were genuine skill failures versus them changing their mind. That distinction is not

Read more
Ships withsolo-founder-skills

Expert skills for non-technical founders building SaaS with AI tools (Claude Code, Lovable, Replit, Cursor). Covers the full lifecycle of planning, building, launching, and growing a software business — actionable guides, checklists, and copy-paste prompts.

Get the whole plugin
Stats
252
Stars
44
Forks
Active
Maintenance
Python
Language
MIT
License
3d ago
Last commit
8mo ago
Created

Repo: whawkinsiv/claude-code-superpowers

Other skills on solo-founder-skills.