agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse,
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/regression-finderContext preview
The summary Claude sees to decide when to auto-load this skill.
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse,
name: regression-finder description: >- Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse, dumber, lazier, or degraded since an update or since they upgraded; asks whether a new Claude Code or Codex version or model made it worse than the old one; wants to know which release, version, or week it regressed in; or wants numbers to report a regression, such as reads before edits, interruptions, and corrections per version. Runs locally and reads transcripts only; nothing goes over the network. license: MIT metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
When a coding agent seems worse after an update, the user's own session history can show whether it changed and when. This skill splits the user's Claude Code or Codex sessions by harness version, model, or week, measures the same behavior in each (reads before the first edit, reads per edit, edits to files not read first, interrupts, corrections, failed tool calls, output, cost), and places the change at the update, or within the span of versions, where the numbers moved, along with anything else that changed at the same point. It reads local transcripts, prints counts and rates, never prints prompt text, and sends nothing anywhere.
Tell the user this when they ask what the check looks at:
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the script path in every command: skill folders can sit under paths with spaces.
1. **Run the default check.** Pick the harness the user asks about (`claude-code` by default, or `codex`) and run:
python3 "<skill-dir>/scripts/regress.py" --harness claude-code
It reads the last 90 days and splits by harness version. Match the flags to the question:
| The user says | Add | |---|---| | "since the last Codex update" | `--harness codex` | | "since I switched to the new model" | `--by model` | | "worse these past few weeks" | `--by week` | | "in my api repo" | `--project <that folder>` | | "since the spring" | `--since 180d` |
It takes a few seconds per thousand sessions and exits 0 even when it finds nothing. Done when the output starts with a bold headline, or you have told the user the exact error.
2. **Answer the user's own question first.** If the user named an update or a time such as last week, find it in the By version table first. If it was not tested, lead with that (for example: 2.1.280 had 7 sessions, and the test needs 20 on each side), then give the headline as background, not as the answer. The notes list every update that was not tested, with its sessions.
Then read the headline. It is one of these kinds:
Done when the user's own update or time is answered, or you know which kind of headline it is.
3. **Follow each confounder.** A confounder is anything else that changed at the same update and could explain the numbers. The section "What else changed at the same point" lists th
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on…