Skip to content
Content
Skill

/sandbox-check

Sandbox audit that probes, from inside a live coding-agent session, what the agent can really reach: secret files it can open, secret-like environment variables, files that git, the shell, the editor, or the harness later run (git hooks, shell startup files, harness settings),

BOOST
From plugin
best-of-agent-harnesses
1.1k10 skills3 agents
Install
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill sandbox-check --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/sandbox-check

Context preview

The summary Claude sees to decide when to auto-load this skill.

Sandbox audit that probes, from inside a live coding-agent session, what the agent can really reach: secret files it can open, secret-like environment variables, files that git, the shell, the editor, or the harness later run (git hooks, shell startup files, harness settings),

SKILL.md

sandbox-check.SKILL.md
name: sandbox-check
description: >-
  Sandbox audit that probes, from inside a live coding-agent session, what the
  agent can really reach: secret files it can open, secret-like environment
  variables, files that git, the shell, the editor, or the harness later run
  (git hooks, shell startup files, harness settings), writes outside the
  project, the Docker socket, the SSH agent, and passwordless sudo. Use when
  the user asks what the agent can access or write on this machine, whether
  the sandbox actually works, how big the blast radius is if the agent gets
  prompt-injected, whether SSH keys, cloud credentials, or .env files are
  exposed, or whether it could escape through Docker or sudo. Runs locally and
  edits no existing file: it creates and deletes one empty test file per
  folder it checks, and only the optional --network flag makes DNS lookups
  and TCP connections.
license: MIT
compatibility: "Python 3.9+ on macOS or Linux; Python 3.11+ to compare Codex settings. Without flags nothing touches the network. The optional --network flag makes DNS lookups and TCP connections (no data sent) to github.com, pypi.org, and the cloud metadata address 169.254.169.254."
metadata:
  author: "Ryan Alberts"
  version: "1.0.0"
  source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"

Sandbox check

A coding agent's shell commands have whatever access that shell has, and a sandbox is only as good as what actually gets through it. This skill runs one probe through the agent's own shell tool and reports, headline first, what the agent can reach right now: secret files, secret-like environment variables, the files other programs later trust and run (a **trust handoff**), the Docker socket, the SSH agent, sudo, and, only when asked, the network. It reads no secret, prints no secret value, and sends nothing anywhere unless you add `--network`.

When to use

  • The user asks what the agent can read, write, or reach on this machine or in this container.
  • The user wants to know whether a sandbox really holds: Claude Code `/sandbox`, Codex sandbox

modes, Gemini CLI `--sandbox`, or Cursor's sandbox.

  • The user worries about prompt injection or a rogue agent and asks for the blast radius.
  • The user asks whether SSH keys, cloud credentials, `.env` files, or tokens are exposed to the

agent.

  • Before an agent works unattended or on an untrusted repository.

When not to use

  • Testing whether permission rules or hooks block dangerous commands: use `guardrail-tester`.
  • Stopping a session that loops or overspends: use `runaway-guard`.
  • Checking which instruction files each agent loads: use `agents-md-checker`.
  • Finding secrets committed to git history: use a secret scanner such as gitleaks. This skill checks

access, not file contents.

  • Turning a sandbox on: that is a settings change. This skill measures the result, and

`references/fixes.md` names the setting.

What the probe touches

Tell the user this, in short, before the first run. It is the whole contract:

  • **Secret files**: opened and closed without reading a byte. Files that macOS keeps only in the

cloud are skipped, so they stay in the cloud.

  • **Existing files other programs trust**: opened for append and closed at once, so contents and

modified times stay the same.

  • **Folders**: one empty file named `.sandbox-check-` plus random letters and `.tmp` is created and

deleted at once. That includes the login-items folder (`Library/LaunchAgents` on macOS; the autostart and systemd user folders on Linux). The folder's own modified time changes, and a file watcher (a dev server, a sync app) may notice for a moment. A test file that cannot be deleted is named in the report.

  • **Targets that do not exist** stay absent: the probe tests only what is there.
  • **Docker socket and SSH agent**: connect, then close, with no data sent. A socket-activated Docker

or Podman service starts when something connects, so the probe can start it.

  • **sudo**: `sudo -n true`, the form that fails instead of asking for a password. The system log

may record the attempt; `--skip-sudo` leaves sudo alone.

  • **Environment variables**: names and value lengths only.
  • **Settings files** of Claude Code, Codex, and Gemini CLI: parsed for their sandbox keys, which

are the only part the report shows.

  • **Network**: used only with `--network`.

On a work machine, endpoint security and audit rules may log or flag the probe: the sudo attempt, the test file in the login-items folder, and the write-opens of shell startup files and the SSH `authorized_keys` file. On Linux, each append check also sends a file-close event to any program watching that file.

`scripts/targets.json` lists every path the probe checks, so it names credential files by design. Each of those lines carries the marker `skillscan:allow`, which tells this repository's security scanner that the path is a probe target, not a file the skill reads.

Steps

`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Run every command from the user's project folder.

1. **Explain the probe** in two or three sentences drawn from "What the probe touches", including the sudo attempt and the login-items test file that security tools may flag. Done when the user says to go ahead. Wait for that answer before step 2.

2. **Run the probe once, through your normal shell tool:**

   python3 "<skill-dir>/scripts/probe.py" --project .

`--project` is the folder the agent works in. When `python3.11` or newer is installed (such as `python3.12` or `python3.13`), use it in place of `python3`: the stock macOS `python3` is 3.9, which skips the Codex settings comparison. The run takes about a second and exits 0 even when checks are blocked: blocked checks are the result, not an error. Run it exactly as sandboxed as every other command in this session, and report what the sandbox blocked as blocked.

Read more
Ships withbest-of-agent-harnesses

🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.

Get the whole plugin

Other skills on best-of-agent-harnesses.