agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap. Use when the user wants a spending cap, budget limit, or cost ceiling for interactive agent sessions (the built-in
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill runaway-guard --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/runaway-guardContext preview
The summary Claude sees to decide when to auto-load this skill.
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap. Use when the user wants a spending cap, budget limit, or cost ceiling for interactive agent sessions (the built-in
name: runaway-guard description: >- Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap. Use when the user wants a spending cap, budget limit, or cost ceiling for interactive agent sessions (the built-in --max-budget-usd works only in print mode); wants to halt an agent that repeats the same command, is stuck in a loop, or runs up a bill unattended overnight; wants a circuit breaker for several failed tool calls in a row; or asks how much the current session has spent against its cap, why a call was blocked, or how to raise or reset the cap. Runs locally: reads the session transcripts before each tool call and sends nothing. license: MIT compatibility: "Python 3.9+ on macOS or Linux, and Claude Code or Codex with hooks enabled. Makes no network calls." metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
An agent left alone can repeat a failing command, retry broken calls, or keep spending long after the work stopped paying off, and Claude Code's `--max-budget-usd` cap works only in print mode. This skill installs a **hook** (a command the harness runs before every tool call) that steps in at three trip wires: the same call a third time with nothing changed, five failed calls in a row, and a dollar cap. The hook reads the session's transcript files on this machine, keeps only counts and hashes, and sends nothing anywhere.
leaving it unattended.
or how to raise the cap or start the count over.
detect loops; `references/harness-hooks.md` says what each lacks.
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command through `python3` exactly as written, with the path in double quotes: skill folders can sit under paths with spaces.
1. **Explain the three trip wires and their defaults** in plain words:
is blocked. Reads and searches in between change nothing; an edit, another command, or a new user prompt does. Re-running tests after an edit never trips it, and neither do deliberate waits or starting subagents.
runs. Codex hooks cannot ask, so there the call is blocked with a message telling the agent to ask the user.
of the cap; at the cap every tool call is blocked (in Claude Code the turn ends) until the user raises the cap or starts the count over. A model missing from the price table is priced like the most expensive known model of its family, so the cap still holds.
Done when the user has heard all three wires and the $10 default cap.
2. **Ask for three choices**: the dollar cap, the harness (Claude Code or Codex), and the scope (user: every session of this user; project: only sessions in the current project). Say that on a subscription plan the dollar figure measures usage at API prices, not the bill. Done when you have all three answers, or the user accepted the defaults: Claude Code, user scope, $10.
3. **Show the dry run** with the user's choices:
python3 "<skill-dir>/scripts/install.py" --harness claude-code --scope user --cap 10
It prints the headline, the settings file it would change, the diff, and notes, and writes nothing. Add `--project <folder>` for project scope when the agent's working folder is not the project. Done when the user has seen the headline and the diff and has said yes or no.
4. **Install only on a clear yes**: run the same command with `--write` added. Done when the output starts with the headline and says "Done". Exit code 2 means a settings file could not be read as JSON, has an unexpected shape, or could not be written; nothing was changed. Quote the message, and leave the file for the user to fix.
5. **Relay the notes** printed under "Notes", in particular:
already counts is blocked at its next tool call once it is over the cap; a session it sees for the first time after spending more than the cap is counted from that point.
from the new place: the hook entry points at this folder.
Done when every note is passed on.
6. **Offer status**:
python3 "<skill-dir>/scripts/status.py"
Done when the user has the headline, or declined.
To remove the guard, run `python3 "<skill-dir>/scripts/install.py" --uninstall`, show the diff, and add `--write` on a clear yes. With no `--harness`, `--scope`, or `--settings`, it checks all four places the guard can b
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…