agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed, and whether code changed after it. Use when the user asks whether the agent really
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill claim-check --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/claim-checkContext preview
The summary Claude sees to decide when to auto-load this skill.
Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed, and whether code changed after it. Use when the user asks whether the agent really
name: claim-check description: >- Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed, and whether code changed after it. Use when the user asks whether the agent really ran the tests, how often it said tests passed without proof, or whether "all tests pass" was true; wants to catch false or stale success claims; asks whether the agent deleted, skipped, or xfailed failing tests or loosened assertions in the current diff; or wants a Stop hook that sends the agent back to rerun tests before finishing. Runs locally: reads Claude Code, Codex, Gemini CLI, and OpenCode transcripts and the git diff; sends nothing. license: MIT compatibility: "Python 3.9+ on macOS or Linux. The diff check needs git. No network access." metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
Coding agents often say "all tests pass" after editing code they never tested again, or after a run that failed. This skill reads the agent's own session transcripts and labels every "tests pass" and "build is clean" claim by the evidence before it, checks the current git diff for weakened tests, and can install a Stop hook that sends the agent back to rerun the tests. It reads transcripts and the repository on this machine; nothing is sent anywhere.
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the path in every command: skill folders can sit under paths with spaces.
1. **Scan the sessions** for the window the user named (default 30 days):
python3 "<skill-dir>/scripts/claims.py" scan --since 30d
Add `--harness claude-code` (or `codex`, `gemini-cli`, `opencode`) or `--project <folder>` when the user asks about one agent or one project, and `--json` when you need every field. Done when the output starts with a bold headline sentence, or you have told the user that no sessions or no claims were found in the window (exit code 0 either way; exit code 2 means a bad argument).
2. **Check the current change** when the user asks about the diff, weakened tests, or deleted tests:
python3 "<skill-dir>/scripts/claims.py" diff --repo .
Use `--base main` (or the branch they name) to check the whole branch from where it left that branch. Done when the output starts with a bold headline, or you have reported the error (exit code 2: not a git repository, or an unknown base).
3. **Offer the Stop hook** only after the report, as its own choice. Show the dry run first:
python3 "<skill-dir>/scripts/install.py" --scope user
Show the user the printed lines and say what the hook does: when the agent tries to finish, it blocks once per reply if the last test run failed, or code changed after the last passing run, in work done since the user's last message. Subagents still working and runs in other repositories do not count. Run the same command with `--write` only after a clear yes. For Codex add `--harness codex`, and tell the user that Codex runs a new hook only after they trust it in `/hooks`. For one project in Claude Code, use `--scope local` (this project, only this user): `--scope project` writes this machine's absolute path to the script into the shared project settings, so use it only when the skill sits inside the repository at the same path for everyone. Done when the user declined, or the output ends with "Added the claim-check Stop hook to ...". The first `--write` keeps the original file as `<file>.claim-check.bak`.
Tell the user to remove the hook before they move, update, or remove this skill: `python3 "<skill-dir>/scripts/install.py" --uninstall --write` (with the same `--harness` and `--scope`). It restores the backup when nothing else in the file changed. The installed command falls back to "allow" if the script is missing, so a moved skill never traps a session, but the stale entry stays in their settings until removed.
Each claim gets one label from the evidence before it in the same session (subagents included):
result could not be read (for example piped through `grep -c`), a later command may have run tests in a way the check cannot read, code changed elsewhere in the repository after a run in one of its subfolders, another subagent changed code after the run, the claim names another command than the last run, a subagent was still working when the claim was made, a subagent's transcript has no event times, the claim names one test while the run had other failures, or the
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on…