Skip to content
Development
Skill

/run-test-plan

Execute YAML test plan, stop on first failure, output rich debug prompt

From plugin
beagle
82139 skills2 commands
Install
$ npx -y skills add existential-birds/beagle --skill run-test-plan --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/run-test-plan

Context preview

The summary Claude sees to decide when to auto-load this skill.

Execute YAML test plan, stop on first failure, output rich debug prompt

SKILL.md

run-test-plan.SKILL.md
description: Execute YAML test plan, stop on first failure, output rich debug prompt
name: run-test-plan
disable-model-invocation: true

Run Test Plan

Execute a YAML test plan, run setup commands, health checks, and each test sequentially. Stop on first failure with rich debug output.

Prerequisites

  • **agent-browser CLI**: Browser tests require the `agent-browser` command-line tool to be available on `PATH`

Arguments

  • `--plan <path>`: Path to test plan (default: `docs/testing/test-plan.yaml`)
  • `--skip-setup`: Skip setup commands and health checks (for re-running after failure)

Step 1: Parse Test Plan

Read and validate the test plan:

# Resolve plan path from --plan (default shown)
PLAN_PATH="${PLAN_PATH:-docs/testing/test-plan.yaml}"

# Check file exists
ls "$PLAN_PATH" || { echo "Error: Test plan not found: $PLAN_PATH"; exit 1; }

# Validate YAML
python3 -c "import yaml; yaml.safe_load(open('$PLAN_PATH'))" || { echo "Error: Invalid YAML: $PLAN_PATH"; exit 1; }

Extract from the YAML:

  • `setup.commands`: List of setup commands
  • `setup.health_checks`: List of URLs to poll
  • `tests`: Array of test cases

Step 2: Run Setup (unless --skip-setup)

2a. Check Prerequisites

If `setup.prerequisites` exists, verify each one:

# For each prerequisite in setup.prerequisites
<prerequisite.check> || { echo "Prerequisite not met: <prerequisite.name>"; exit 1; }

2b. Set Environment Variables

If `setup.env` exists, export each variable. Variables using `${VAR}` syntax should be resolved from the current environment:

# For each key/value in setup.env
export <key>="<value>"

2c. Build

If `setup.build` exists, execute build commands sequentially:

# For each command in setup.build
<command> || { echo "Build failed: <command>"; exit 1; }

2d. Start Services

If `setup.services` exists, start long-running processes and wait for health checks:

# For each service in setup.services
nohup <service.command> > .beagle/service-<index>.log 2>&1 &
echo $! > .beagle/service-<index>.pid

For each service with a `health_check`, poll until ready:

timeout=<service.health_check.timeout or 30>
url=<service.health_check.url>
elapsed=0

while [ $elapsed -lt $timeout ]; do
  if curl -s -o /dev/null -w "%{http_code}" "$url" | grep -qE "^(200|301|302)"; then
    echo "✓ Health check passed: $url"
    break
  fi
  sleep 2
  elapsed=$((elapsed + 2))
done

if [ $elapsed -ge $timeout ]; then
  echo "✗ Health check timeout: $url"
  exit 1
fi

2e. Legacy Setup Format

If the plan uses the older flat format (`setup.commands` + `setup.health_checks` instead of `prerequisites`/`build`/`services`), fall back to executing `setup.commands` sequentially and polling `setup.health_checks` as before.

Step 3: Gate — setup ready before tests

Do not start **Step 4** until each condition you can check is true:

1. **Plan load:** The file from `--plan` (default `docs/testing/test-plan.yaml`) exists and parses as YAML (same checks as Step 1). 2. **Setup branch:**

  • If **not** using `--skip-setup`: Every `setup.prerequisites` check that exists exited 0; every `setup.build` command succeeded; every service `health_check` reached HTTP 200, 301, or 302 within its timeout **or** legacy `setup.health_checks` passed after `setup.commands`.
  • If using `--skip-setup`: Before TC-01, confirm anything the plan still needs is alive—at minimum one successful `curl` (or equivalent) to each URL in `setup.health_checks` or each `setup.services[].health_check.url` that the tests depend on.

3. **Evidence path:** `mkdir -p docs/testing/evidence` succeeds and the directory exists.

If any gate fails, stop, fix setup or flags, and do not execute tests.

Step 4: Execute Tests Sequentially

For each test in the plan:

4a. Log Test Start

## Running: TC-XX - <test.name>

Context: <test.context>

4b. Execute Steps

For each step in `test.steps`, determine the step type and execute accordingly:

**Shell commands (`run:` steps):**

The most common step type. Execute the command via Bash and capture stdout, stderr, and exit code:

# Execute the command, capture output and exit code
<command> 2>&1
echo "EXIT_CODE: $?"

Capture all output for evaluation in step 4c. Shell steps cover:

  • CLI binary invocations (e.g., `./target/debug/myapp status --all`)
  • Database queries (e.g., `psql "${DATABASE_URL}" -c "SELECT ..."`)
  • File inspection (e.g., `ls -la /path/to/expected/output`)
  • Process lifecycle checks (e.g., `timeout 5 ./myapp 2>&1 || true`)
  • Any other command a human would type in a terminal

**curl actions (`action: curl` steps):**

curl -X <method> \
  -H "Content-Type: application/json" \
  <additional headers> \
  -d '<body>' \
  "<url>" \
  -o response.json \
  -w "%{http_code}" > status_code.txt

# Capture response for evaluation
cat response.json
cat status_code.txt

**agent-browser CLI actions:**

Steps starting with `agent-browser` are browser automation commands:

# Navigate
agent-browser open <url>

# Snapshot interactive elements (always do before interacting)
agent-browser snapshot -i

# Interact using refs from snapshot output (@e1, @e2, etc.)
agent-browser fill @<ref> "<value>"
agent-browser click @<ref>

# Wait for conditions
agent-browser wait --url "<pattern>"
agent-browser wait --text "<text>"
agent-browser wait --load networkidle

# Capture evidence
agent-browser screenshot docs/testing/evidence/<test.id>.png

**Important:** Always run `agent-browser snapshot -i` before interacting with elements to get valid refs, and re-snapshot after navigation or significant DOM changes.

Save screenshots to `docs/testing/evidence/<test.id>.png`

4c. Evaluate Result

**Gate — artifacts before PASS/FAIL:**

  • **`run:` steps:** Stdout and stderr captured; exit code recorded (e.g. `EXIT_CODE:` line or equivalent).
  • **`action: curl` steps:** Response body and HTTP statu
Read more
Ships withbeagle

Image: NASA, Public Domain. Source Beagle is an Agent Skills marketplace: framework-aware code review, documentation, testing, architectural analysis, and git workflows for any compatible coding agent.

Get the whole plugin

Other skills on beagle.