agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/guardrail-testerContext preview
The summary Claude sees to decide when to auto-load this skill.
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures
name: guardrail-tester description: >- Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures prompt friction on the latest tool calls. Use when the user asks whether deny rules or hooks block force pushes, rm -rf, secret reads, or downloads piped into a shell; wants to test or audit guardrails, or find a bypass in permission settings or exec policy; asks how many prompts or blocks the rules cause; or just installed a guard hook and wants proof it holds. Local only, no network; executes the user's own hook commands only with their yes. license: MIT compatibility: "Python 3.9+, standard library only. Reading Codex config.toml and Gemini CLI policy files needs Python 3.11+. Makes no network calls; the only commands it runs are the user's own hook commands, fed test JSON, and only with --run-hooks after the user says yes." metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
Permission rules match the text of a command, so `git push origin +main`, `rm -r -f build`, and `bash -c '...'` can walk past a deny rule written for the plain form. This skill checks the rules and hooks the user already has against about 90 dangerous commands in those forms, then replays their recent tool calls to count how often the rules would ask or block. The user gets a headline ("Your Claude Code guardrails block 58 of 90 dangerous commands outright. 24 more stop at a prompt..."), every miss with a tested fix, and their friction numbers. Everything is read and simulated locally; nothing is sent anywhere, and it runs the user's hook commands only after they say yes.
repository, then test it here.
Say this to the user in short before the first run with hooks. It is the whole contract.
The tester never runs them; a hook that runs or forwards its input would. Each battery command is also written to do nothing if run: its targets sit under a missing `./guardrail-tester-probe/` folder, its remotes and branches (`probe-remote`, `probe-branch`) are made up, and its hosts end in `.invalid`, a name reserved for addresses that resolve nowhere.
"simulated". `references/harness-rules.md` says what each harness simulation covers.
PreToolUse hook command then runs once per battery case, one run at a time, with a 10-second limit, from the folder its harness uses (usually the project), with the session's environment variables plus `CLAUDE_PROJECT_DIR`. A hook that writes a log or keeps state records those test calls. A hook that times out twice is not run again.
and checks the rules only. With `--replay-hooks` (which needs `--run-hooks`) it also runs this project's hooks on replayed calls from this project. Calls from other project folders are checked against their own rules; their hooks never run.
`scripts/battery.json` holds the dangerous commands on purpose: they are the test. Each line carries the marker `skillscan:allow`, which tells this repository's security scanner that the line is test data, not a command the skill runs.
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the user's project folder, and give each Bash call a 10-minute timeout (600000 ms): hooks run one at a time, so a slow hook makes the run long.
1. **Run the rules-only test**, which reads settings and runs no hook:
python3 "<skill-dir>/scripts/test_guards.py" --project . --harness claude-code
Add `--harness` for the harness you are running in (`claude-code`, `codex`, `gemini-cli`, `opencode`, or `cursor`) unless the user asks about all; without it the tester covers every harness with settings on this machine. Done when the output starts with a bold headline, or you have told the user the error (exit code 2 means a bad argument or an unreadable battery file).
2. **Tell the user what was found**: the settings files and every hook command from the report's "What was found" section, quoted exactly. Done when the user has seen the list.
3. **Read each hook script, then ask before running hooks.** Open the script each hook command runs. If a script runs, evaluates, or sends its input anywhere (`eval`, `bash -c "$cmd"`, a `curl` with the input, a queue), or has other side effects such as writing a log, tell the user exactly that and recommend testing without hooks. Otherwise name the hook commands and say each will get test JSON for about 90 battery calls. On a clear yes, run this with a 10-minute Bash time
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on…