agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size.
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/harness-test-driveContext preview
The summary Claude sees to decide when to auto-load this skill.
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size.
name: harness-test-drive description: >- Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size. Use when the user asks which coding agent or harness works best on their codebase; wants to compare or benchmark agents on their own repository instead of trusting leaderboards such as SWE-bench; wants to trial one agent against another before switching; or asks what each agent costs per fixed bug. Spends money: the harness CLIs call their model providers, so a dollar cap is required. The scripts make no network calls. license: MIT compatibility: "Python 3.9+ and git on macOS or Linux, plus the harness CLIs to compare (claude, codex, gemini). The scripts make no network calls themselves; each harness CLI sends the prompt and the code it reads to its model provider over the network, and those runs cost money. The repository's test command must run on this machine." metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
Public leaderboards measure someone else's code, and harness rankings barely carry over from one repository to the next. This skill runs coding agents on tasks taken from the user's own git history: past commits whose tests failed before the change and passed after it. Each agent gets the commit message as its prompt in a fresh copy of the repository, and the repository's own tests decide whether it passed. The result is a scoreboard of passes, dollars per pass, minutes, and change size. The scripts read the repository and send nothing anywhere; each harness sends the prompt and the code it reads to its model provider, as it does in normal use, and that costs money.
switching.
the scores come from the tests, and point to the manual method in [How to test-drive a harness](https://github.com/RyanAlberts/best-of-Agent-Harnesses/blob/main/comparisons/how-to-test-drive-a-harness.md).
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the root of the user's repository, with the skill path in quotes as shown. Inside the repository the scripts write only into `.harness-test-drive/`, a folder that ignores itself in git. Every check and every agent run works in its own temporary copy, deleted afterwards.
**Before step 2, for a repository the user did not write, recommend a container or VM.** Step 2 runs the repository's test and setup commands at up to 31 commits on this computer. In step 5 each harness loads the repository's own agent settings (Gemini CLI with the copy trusted), edits files, runs the test command, and reads commit messages as prompts.
1. **Find the test command and the candidate tasks.** This reads git history and runs nothing:
python3 "<skill-dir>/scripts/mine_tasks.py" --repo .
It prints the test command it detected, or exits 2 asking for one. Confirm the command with the user. Each check runs in a fresh copy with nothing installed, so when the tests need dependencies, add a `--setup-cmd` that installs inside the copy: `npm ci`, `pnpm install --frozen-lockfile`, or `python3 -m venv .venv && .venv/bin/pip install -e .` with the test command `.venv/bin/python -m pytest`. Never install into the user's own environment. Prefer a test command that runs offline and writes only inside the copy, so Codex's sandbox can run it too. Done when the user has confirmed a test command and the report shows at least one candidate, or you have told the user why there is none.
2. **Check the tasks.** For each candidate, in a fresh copy of the repository, the test command must fail at the commit before the change (with its tests added) and pass with the whole change, twice:
python3 "<skill-dir>/scripts/mine_tasks.py" --repo . --test-cmd "<test-cmd>" --validate
Add the `--setup-cmd` from step 1 when there is one. It first checks that the tests pass at HEAD, then keeps up to 10 tasks (`--max`). Expect two to four test-suite runs per candidate and up to 30 candidates, so run it in the background when the suite takes more than a minute. Done when the headline says how many tasks were kept. Exit 2 with "fails at HEAD" means the command or the setup is wrong. The message ends with the test output. A missing module or command means the clean copy needs `--setup-cmd`. Fix it with the user and rerun. Fewer than five tasks makes a weak comparison; say so, and offer `--since 730d` for more history.
3. **Choose the harnesses and show the estimate:**
python3 "<skill-dir>/scripts/drive.py" estimate --tasks .harness-test-drive/tasks.json
It lists which harnesses are on this computer, the number of runs, and a dollar range per harness. Ask the user to pick two or three. Done when the user has picked the harnesses and seen the range for them (rerun with `--harness` to show only those).
4. **Get an explicit dollar cap.** This is the one skill in the set that spends money. As
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on…