Skip to content
Content
Skill

/regression-finder

Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse,

BOOST
From plugin
best-of-agent-harnesses
1.1k10 skills3 agents
Install
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/regression-finder

Context preview

The summary Claude sees to decide when to auto-load this skill.

Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse,

SKILL.md

regression-finder.SKILL.md
name: regression-finder
description: >-
  Regression check for coding agents: shows how the agent behaved before and
  after each harness update, model switch, or week in the user's own Claude
  Code or Codex history, and finds the point where it changed. Use when the
  user says the agent's quality dropped or it got worse, dumber, lazier, or
  degraded since an update or since they upgraded; asks whether a new Claude
  Code or Codex version or model made it worse than the old one; wants to know
  which release, version, or week it regressed in; or wants numbers to report a
  regression, such as reads before edits, interruptions, and corrections per
  version. Runs locally and reads transcripts only; nothing goes over the
  network.
license: MIT
metadata:
  author: "Ryan Alberts"
  version: "1.0.0"
  source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"

Regression finder

When a coding agent seems worse after an update, the user's own session history can show whether it changed and when. This skill splits the user's Claude Code or Codex sessions by harness version, model, or week, measures the same behavior in each (reads before the first edit, reads per edit, edits to files not read first, interrupts, corrections, failed tool calls, output, cost), and places the change at the update, or within the span of versions, where the numbers moved, along with anything else that changed at the same point. It reads local transcripts, prints counts and rates, never prints prompt text, and sends nothing anywhere.

When to use

  • The user says the agent got worse, dumber, or lazier after an update, or asks "is it just me?"
  • The user asks whether a new Claude Code or Codex version, or a new model, changed how the agent works.
  • The user wants to know which release or week a regression started.
  • The user wants evidence for a bug report about a regression, like [anthropics/claude-code#42796](https://github.com/anthropics/claude-code/issues/42796).

When not to use

  • Where tokens and money go in general (re-reads, cache rebuilds, oversized results): use `session-waste-report`.
  • Stopping a session that loops or overspends right now: use `runaway-guard`.
  • Comparing two harnesses on the same tasks: use `harness-test-drive`.
  • Checking whether "tests pass" claims were true: use `claim-check`.
  • Finding which instruction-file rules the agent breaks: use `rules-to-guards`.
  • Cursor: it keeps no usable transcripts, so there is nothing to measure.

What it reads

Tell the user this when they ask what the check looks at:

  • **Read**: the session files each harness keeps (Claude Code under `~/.claude/projects/`, Codex under `~/.codex/sessions/`), only those changed within the window. Prompt text is read on this machine only to spot corrections such as "no, that's wrong".
  • **Printed**: counts, rates, token totals, versions, model names, and project folders. Never prompt text, commands, or file contents.
  • **Written**: nothing, unless the user asks for `--out` or `--svg`, which write the one file named.
  • **Sent**: nothing. It makes no network calls.

Steps

`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the script path in every command: skill folders can sit under paths with spaces.

1. **Run the default check.** Pick the harness the user asks about (`claude-code` by default, or `codex`) and run:

   python3 "<skill-dir>/scripts/regress.py" --harness claude-code

It reads the last 90 days and splits by harness version. Match the flags to the question:

| The user says | Add | |---|---| | "since the last Codex update" | `--harness codex` | | "since I switched to the new model" | `--by model` | | "worse these past few weeks" | `--by week` | | "in my api repo" | `--project <that folder>` | | "since the spring" | `--since 180d` |

It takes a few seconds per thousand sessions and exits 0 even when it finds nothing. Done when the output starts with a bold headline, or you have told the user the exact error.

2. **Answer the user's own question first.** If the user named an update or a time such as last week, find it in the By version table first. If it was not tested, lead with that (for example: 2.1.280 had 7 sessions, and the test needs 20 on each side), then give the headline as background, not as the answer. The notes list every update that was not tested, with its sessions.

Then read the headline. It is one of these kinds:

  • A flagged change: "After Claude Code `2.1.270`, your agent reads 41% less before it edits...", or "After an update between Claude Code `2.1.260` and `2.1.270`, ..." when the test had to borrow neighboring versions. The change lies somewhere in that span; say the span, not one version. Go to step 3.
  • "No behavior change passed the test across ...": the numbers wobble but nothing passed. Say so plainly and give the session counts; go to step 5.
  • "No lasting behavior change passed the test ...; 1 version stands out": one version differs from the versions on both sides of it. Report the "Stands out" line as it is.
  • "Not enough history to test an update yet": fewer than 20 sessions on each side of every update. Offer, in this order, a longer window (`--since 180d`, when the history goes back that far), a split by time (`--by week`), and last a lower bar (`--min-sessions 12`, the floor). A lower bar tests more updates, but each test can only catch larger changes.
  • "No ... found", "records no version", "ran on one ...", or "Not enough history to compare models": nothing to compare yet. Say so, and name what the notes list as left out.

Done when the user's own update or time is answered, or you know which kind of headline it is.

3. **Follow each confounder.** A confounder is anything else that changed at the same update and could explain the numbers. The section "What else changed at the same point" lists th

Read more
Ships withbest-of-agent-harnesses

🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.

Get the whole plugin

Other skills on best-of-agent-harnesses.