Skip to content
Content
Skill

/claim-check

Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed, and whether code changed after it. Use when the user asks whether the agent really

BOOST
From plugin
best-of-agent-harnesses
1.1k10 skills3 agents
Install
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill claim-check --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/claim-check

Context preview

The summary Claude sees to decide when to auto-load this skill.

Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed, and whether code changed after it. Use when the user asks whether the agent really

SKILL.md

claim-check.SKILL.md
name: claim-check
description: >-
  Claim checker that audits a coding agent's statements that tests pass or a
  build is clean against its own session transcripts: whether a matching run
  happened before the claim, whether it passed, and whether code changed
  after it. Use when the user asks whether the agent really ran the tests,
  how often it said tests passed without proof, or whether "all tests pass"
  was true; wants to catch false or stale success claims; asks whether the
  agent deleted, skipped, or xfailed failing tests or loosened assertions in
  the current diff; or wants a Stop hook that sends the agent back to rerun
  tests before finishing. Runs locally: reads Claude Code, Codex, Gemini CLI,
  and OpenCode transcripts and the git diff; sends nothing.
license: MIT
compatibility: "Python 3.9+ on macOS or Linux. The diff check needs git. No network access."
metadata:
  author: "Ryan Alberts"
  version: "1.0.0"
  source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"

Claim check

Coding agents often say "all tests pass" after editing code they never tested again, or after a run that failed. This skill reads the agent's own session transcripts and labels every "tests pass" and "build is clean" claim by the evidence before it, checks the current git diff for weakened tests, and can install a Stop hook that sends the agent back to rerun the tests. It reads transcripts and the repository on this machine; nothing is sent anywhere.

When to use

  • The user asks whether the agent really ran the tests, or how often it claimed success without proof.
  • The user doubts a specific "tests pass" or "build succeeds" message.
  • The user asks whether the current change deleted, skipped, or loosened tests.
  • The user wants the agent stopped from finishing while the last test run failed or went stale.

When not to use

  • Comparing which agent fixes bugs best: use `harness-test-drive`.
  • Making the agent obey a written rule such as "run tests before committing": use `rules-to-guards`.
  • Blocking dangerous commands: use `guardrail-tester`.
  • Loops and spending caps: use `runaway-guard`. Token and money waste: use `session-waste-report`.
  • Whether the test commands in AGENTS.md still work: use `agents-md-checker`.

Steps

`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the path in every command: skill folders can sit under paths with spaces.

1. **Scan the sessions** for the window the user named (default 30 days):

   python3 "<skill-dir>/scripts/claims.py" scan --since 30d

Add `--harness claude-code` (or `codex`, `gemini-cli`, `opencode`) or `--project <folder>` when the user asks about one agent or one project, and `--json` when you need every field. Done when the output starts with a bold headline sentence, or you have told the user that no sessions or no claims were found in the window (exit code 0 either way; exit code 2 means a bad argument).

2. **Check the current change** when the user asks about the diff, weakened tests, or deleted tests:

   python3 "<skill-dir>/scripts/claims.py" diff --repo .

Use `--base main` (or the branch they name) to check the whole branch from where it left that branch. Done when the output starts with a bold headline, or you have reported the error (exit code 2: not a git repository, or an unknown base).

3. **Offer the Stop hook** only after the report, as its own choice. Show the dry run first:

   python3 "<skill-dir>/scripts/install.py" --scope user

Show the user the printed lines and say what the hook does: when the agent tries to finish, it blocks once per reply if the last test run failed, or code changed after the last passing run, in work done since the user's last message. Subagents still working and runs in other repositories do not count. Run the same command with `--write` only after a clear yes. For Codex add `--harness codex`, and tell the user that Codex runs a new hook only after they trust it in `/hooks`. For one project in Claude Code, use `--scope local` (this project, only this user): `--scope project` writes this machine's absolute path to the script into the shared project settings, so use it only when the skill sits inside the repository at the same path for everyone. Done when the user declined, or the output ends with "Added the claim-check Stop hook to ...". The first `--write` keeps the original file as `<file>.claim-check.bak`.

Tell the user to remove the hook before they move, update, or remove this skill: `python3 "<skill-dir>/scripts/install.py" --uninstall --write` (with the same `--harness` and `--scope`). It restores the backup when nothing else in the file changed. The installed command falls back to "allow" if the script is missing, so a moved skill never traps a session, but the stale entry stays in their settings until removed.

Read the results

Each claim gets one label from the evidence before it in the same session (subagents included):

  • **backed**: the latest matching run passed, and no code changed after it.
  • **stale**: the run passed, but code changed after it and nothing ran again.
  • **contradicted**: the latest matching run failed.
  • **unsupported**: no matching run happened in the session before the claim.
  • **unclear**: the evidence could not be read or ordered. The `why` field says which case: the

result could not be read (for example piped through `grep -c`), a later command may have run tests in a way the check cannot read, code changed elsewhere in the repository after a run in one of its subfolders, another subagent changed code after the run, the claim names another command than the last run, a subagent was still working when the claim was made, a subagent's transcript has no event times, the claim names one test while the run had other failures, or the

Read more
Ships withbest-of-agent-harnesses

🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.

Get the whole plugin

Other skills on best-of-agent-harnesses.