Skip to content
Content
Skill

/harness-test-drive

Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size.

BOOST
From plugin
best-of-agent-harnesses
1.1k10 skills3 agents
Install
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill harness-test-drive --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/harness-test-drive

Context preview

The summary Claude sees to decide when to auto-load this skill.

Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests score how many tasks it completes, dollars per task, minutes, and change size.

SKILL.md

harness-test-drive.SKILL.md
name: harness-test-drive
description: >-
  Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined
  from the user's own git history: each agent gets a past commit message in a
  fresh copy of the repo, and the repo's own tests score how many tasks it
  completes, dollars per task, minutes, and change size. Use when the user asks
  which coding agent or harness works best on their codebase; wants to compare
  or benchmark agents on their own repository instead of trusting leaderboards
  such as SWE-bench; wants to trial one agent against another before
  switching; or asks what each agent costs per fixed bug. Spends money: the
  harness CLIs call their model providers, so a dollar cap is required. The
  scripts make no network calls.
license: MIT
compatibility: "Python 3.9+ and git on macOS or Linux, plus the harness CLIs to compare (claude, codex, gemini). The scripts make no network calls themselves; each harness CLI sends the prompt and the code it reads to its model provider over the network, and those runs cost money. The repository's test command must run on this machine."
metadata:
  author: "Ryan Alberts"
  version: "1.0.0"
  source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"

Harness test drive

Public leaderboards measure someone else's code, and harness rankings barely carry over from one repository to the next. This skill runs coding agents on tasks taken from the user's own git history: past commits whose tests failed before the change and passed after it. Each agent gets the commit message as its prompt in a fresh copy of the repository, and the repository's own tests decide whether it passed. The result is a scoreboard of passes, dollars per pass, minutes, and change size. The scripts read the repository and send nothing anywhere; each harness sends the prompt and the code it reads to its model provider, as it does in normal use, and that costs money.

When to use

  • The user asks which coding agent or harness is best for their own code.
  • The user wants to compare Claude Code, Codex, or Gemini CLI on real tasks before choosing or

switching.

  • The user asks what an agent costs per fixed bug, or how long it takes, on their repository.

When not to use

  • Totals of past spend or wasted tokens: use `session-waste-report`.
  • Whether an agent got worse after an update: use `regression-finder`.
  • Stopping a live session that loops or overspends: use `runaway-guard`.
  • Checking whether an agent's "tests pass" claims were true: use `claim-check`.
  • A repository with no test command that passes, or no commits that change tests: explain that

the scores come from the tests, and point to the manual method in [How to test-drive a harness](https://github.com/RyanAlberts/best-of-Agent-Harnesses/blob/main/comparisons/how-to-test-drive-a-harness.md).

Steps

`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the root of the user's repository, with the skill path in quotes as shown. Inside the repository the scripts write only into `.harness-test-drive/`, a folder that ignores itself in git. Every check and every agent run works in its own temporary copy, deleted afterwards.

**Before step 2, for a repository the user did not write, recommend a container or VM.** Step 2 runs the repository's test and setup commands at up to 31 commits on this computer. In step 5 each harness loads the repository's own agent settings (Gemini CLI with the copy trusted), edits files, runs the test command, and reads commit messages as prompts.

1. **Find the test command and the candidate tasks.** This reads git history and runs nothing:

   python3 "<skill-dir>/scripts/mine_tasks.py" --repo .

It prints the test command it detected, or exits 2 asking for one. Confirm the command with the user. Each check runs in a fresh copy with nothing installed, so when the tests need dependencies, add a `--setup-cmd` that installs inside the copy: `npm ci`, `pnpm install --frozen-lockfile`, or `python3 -m venv .venv && .venv/bin/pip install -e .` with the test command `.venv/bin/python -m pytest`. Never install into the user's own environment. Prefer a test command that runs offline and writes only inside the copy, so Codex's sandbox can run it too. Done when the user has confirmed a test command and the report shows at least one candidate, or you have told the user why there is none.

2. **Check the tasks.** For each candidate, in a fresh copy of the repository, the test command must fail at the commit before the change (with its tests added) and pass with the whole change, twice:

   python3 "<skill-dir>/scripts/mine_tasks.py" --repo . --test-cmd "<test-cmd>" --validate

Add the `--setup-cmd` from step 1 when there is one. It first checks that the tests pass at HEAD, then keeps up to 10 tasks (`--max`). Expect two to four test-suite runs per candidate and up to 30 candidates, so run it in the background when the suite takes more than a minute. Done when the headline says how many tasks were kept. Exit 2 with "fails at HEAD" means the command or the setup is wrong. The message ends with the test output. A missing module or command means the clean copy needs `--setup-cmd`. Fix it with the user and rerun. Fewer than five tasks makes a weak comparison; say so, and offer `--since 730d` for more history.

3. **Choose the harnesses and show the estimate:**

   python3 "<skill-dir>/scripts/drive.py" estimate --tasks .harness-test-drive/tasks.json

It lists which harnesses are on this computer, the number of runs, and a dollar range per harness. Ask the user to pick two or three. Done when the user has picked the harnesses and seen the range for them (rerun with `--harness` to show only those).

4. **Get an explicit dollar cap.** This is the one skill in the set that spends money. As

Read more
Ships withbest-of-agent-harnesses

🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.

Get the whole plugin

Other skills on best-of-agent-harnesses.