Skip to content
Content
Skill

/guardrail-tester

Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures

BOOST
From plugin
best-of-agent-harnesses
1.1k10 skills3 agents
Install
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/guardrail-tester

Context preview

The summary Claude sees to decide when to auto-load this skill.

Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures

SKILL.md

guardrail-tester.SKILL.md
name: guardrail-tester
description: >-
  Guardrail tester that checks whether the permission rules and PreToolUse
  hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor
  stop a battery of dangerous commands, including wrapped, reordered, and
  full-path forms that slip past prefix rules, and measures prompt friction
  on the latest tool calls. Use when the user asks whether deny rules or
  hooks block force pushes, rm -rf, secret reads, or downloads piped into a
  shell; wants to test or audit guardrails, or find a bypass in permission
  settings or exec policy; asks how many prompts or blocks the rules cause;
  or just installed a guard hook and wants proof it holds. Local only, no
  network; executes the user's own hook commands only with their yes.
license: MIT
compatibility: "Python 3.9+, standard library only. Reading Codex config.toml and Gemini CLI policy files needs Python 3.11+. Makes no network calls; the only commands it runs are the user's own hook commands, fed test JSON, and only with --run-hooks after the user says yes."
metadata:
  author: "Ryan Alberts"
  version: "1.0.0"
  source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"

Guardrail tester

Permission rules match the text of a command, so `git push origin +main`, `rm -r -f build`, and `bash -c '...'` can walk past a deny rule written for the plain form. This skill checks the rules and hooks the user already has against about 90 dangerous commands in those forms, then replays their recent tool calls to count how often the rules would ask or block. The user gets a headline ("Your Claude Code guardrails block 58 of 90 dangerous commands outright. 24 more stop at a prompt..."), every miss with a tested fix, and their friction numbers. Everything is read and simulated locally; nothing is sent anywhere, and it runs the user's hook commands only after they say yes.

When to use

  • The user asks whether their deny rules, ask rules, or hooks really stop a dangerous command.
  • The user wants to test, audit, or find bypasses in their guardrails or permission settings.
  • The user asks how many prompts their rules cause, or wants fewer without losing safety.
  • The user just installed a guard (a hook, dcg, the repo's safe-settings template) and wants proof.

When not to use

  • What the agent can reach on this machine (secret files, Docker, sudo): use `sandbox-check`.
  • Turning a rule from AGENTS.md or CLAUDE.md into a hook: use `rules-to-guards`.
  • Stopping a session that loops or overspends: use `runaway-guard`.
  • Building a guard from scratch: recommend dcg or `templates/claude-code-safe-settings/` in this

repository, then test it here.

What the tester runs

Say this to the user in short before the first run with hooks. It is the whole contract.

  • **The battery commands never run.** Each is only text inside the JSON a hook reads on stdin.

The tester never runs them; a hook that runs or forwards its input would. Each battery command is also written to do nothing if run: its targets sit under a missing `./guardrail-tester-probe/` folder, its remotes and branches (`probe-remote`, `probe-branch`) are made up, and its hosts end in `.invalid`, a name reserved for addresses that resolve nowhere.

  • **Rules are simulated** from each harness's documented matching rules, so results are labeled

"simulated". `references/harness-rules.md` says what each harness simulation covers.

  • **It runs your hook commands only after you say yes** (the `--run-hooks` flag). Each matching

PreToolUse hook command then runs once per battery case, one run at a time, with a 10-second limit, from the folder its harness uses (usually the project), with the session's environment variables plus `CLAUDE_PROJECT_DIR`. A hook that writes a log or keeps state records those test calls. A hook that times out twice is not run again.

  • **Replay** reads session transcripts on this machine, read-only, masks secrets in its output,

and checks the rules only. With `--replay-hooks` (which needs `--run-hooks`) it also runs this project's hooks on replayed calls from this project. Calls from other project folders are checked against their own rules; their hooks never run.

`scripts/battery.json` holds the dangerous commands on purpose: they are the test. Each line carries the marker `skillscan:allow`, which tells this repository's security scanner that the line is test data, not a command the skill runs.

Steps

`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the user's project folder, and give each Bash call a 10-minute timeout (600000 ms): hooks run one at a time, so a slow hook makes the run long.

1. **Run the rules-only test**, which reads settings and runs no hook:

   python3 "<skill-dir>/scripts/test_guards.py" --project . --harness claude-code

Add `--harness` for the harness you are running in (`claude-code`, `codex`, `gemini-cli`, `opencode`, or `cursor`) unless the user asks about all; without it the tester covers every harness with settings on this machine. Done when the output starts with a bold headline, or you have told the user the error (exit code 2 means a bad argument or an unreadable battery file).

2. **Tell the user what was found**: the settings files and every hook command from the report's "What was found" section, quoted exactly. Done when the user has seen the list.

3. **Read each hook script, then ask before running hooks.** Open the script each hook command runs. If a script runs, evaluates, or sends its input anywhere (`eval`, `bash -c "$cmd"`, a `curl` with the input, a queue), or has other side effects such as writing a log, tell the user exactly that and recommend testing without hooks. Otherwise name the hook commands and say each will get test JSON for about 90 battery calls. On a clear yes, run this with a 10-minute Bash time

Read more
Ships withbest-of-agent-harnesses

🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.

Get the whole plugin

Other skills on best-of-agent-harnesses.