agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Sandbox audit that probes, from inside a live coding-agent session, what the agent can really reach: secret files it can open, secret-like environment variables, files that git, the shell, the editor, or the harness later run (git hooks, shell startup files, harness settings),
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill sandbox-check --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/sandbox-checkContext preview
The summary Claude sees to decide when to auto-load this skill.
Sandbox audit that probes, from inside a live coding-agent session, what the agent can really reach: secret files it can open, secret-like environment variables, files that git, the shell, the editor, or the harness later run (git hooks, shell startup files, harness settings),
name: sandbox-check description: >- Sandbox audit that probes, from inside a live coding-agent session, what the agent can really reach: secret files it can open, secret-like environment variables, files that git, the shell, the editor, or the harness later run (git hooks, shell startup files, harness settings), writes outside the project, the Docker socket, the SSH agent, and passwordless sudo. Use when the user asks what the agent can access or write on this machine, whether the sandbox actually works, how big the blast radius is if the agent gets prompt-injected, whether SSH keys, cloud credentials, or .env files are exposed, or whether it could escape through Docker or sudo. Runs locally and edits no existing file: it creates and deletes one empty test file per folder it checks, and only the optional --network flag makes DNS lookups and TCP connections. license: MIT compatibility: "Python 3.9+ on macOS or Linux; Python 3.11+ to compare Codex settings. Without flags nothing touches the network. The optional --network flag makes DNS lookups and TCP connections (no data sent) to github.com, pypi.org, and the cloud metadata address 169.254.169.254." metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
A coding agent's shell commands have whatever access that shell has, and a sandbox is only as good as what actually gets through it. This skill runs one probe through the agent's own shell tool and reports, headline first, what the agent can reach right now: secret files, secret-like environment variables, the files other programs later trust and run (a **trust handoff**), the Docker socket, the SSH agent, sudo, and, only when asked, the network. It reads no secret, prints no secret value, and sends nothing anywhere unless you add `--network`.
modes, Gemini CLI `--sandbox`, or Cursor's sandbox.
agent.
access, not file contents.
`references/fixes.md` names the setting.
Tell the user this, in short, before the first run. It is the whole contract:
cloud are skipped, so they stay in the cloud.
modified times stay the same.
deleted at once. That includes the login-items folder (`Library/LaunchAgents` on macOS; the autostart and systemd user folders on Linux). The folder's own modified time changes, and a file watcher (a dev server, a sync app) may notice for a moment. A test file that cannot be deleted is named in the report.
or Podman service starts when something connects, so the probe can start it.
may record the attempt; `--skip-sudo` leaves sudo alone.
are the only part the report shows.
On a work machine, endpoint security and audit rules may log or flag the probe: the sudo attempt, the test file in the login-items folder, and the write-opens of shell startup files and the SSH `authorized_keys` file. On Linux, each append check also sends a file-close event to any program watching that file.
`scripts/targets.json` lists every path the probe checks, so it names credential files by design. Each of those lines carries the marker `skillscan:allow`, which tells this repository's security scanner that the path is a probe target, not a file the skill reads.
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the user's project folder.
1. **Explain the probe** in two or three sentences drawn from "What the probe touches", including the sudo attempt and the login-items test file that security tools may flag. Done when the user says to go ahead. Wait for that answer before step 2.
2. **Run the probe once, through your normal shell tool:**
python3 "<skill-dir>/scripts/probe.py" --project .
`--project` is the folder the agent works in. When `python3.11` or newer is installed (such as `python3.12` or `python3.13`), use it in place of `python3`: the stock macOS `python3` is 3.9, which skips the Codex settings comparison. The run takes about a second and exits 0 even when checks are blocked: blocked checks are the result, not an error. Run it exactly as sandboxed as every other command in this session, and report what the sandbox blocked as blocked.
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…