Skip to content
Development
Skill

/test-run-rules

Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run

BOOST
From plugin
turbo
40573 skills
Install
$ npx -y skills add tobihagemann/turbo --skill test-run-rules --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/test-run-rules

Context preview

The summary Claude sees to decide when to auto-load this skill.

Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run

SKILL.md

test-run-rules.SKILL.md
name: test-run-rules
description: "Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run created. Not typically invoked directly."

Test Run Rules

Apply these rules for the rest of the test run, whether it drives the running app or runs a test target. A scenario they leave **blocked** could not be set up or run; one they leave **inconclusive** ran, but its result cannot be read.

Sandbox Denials

When a command failed on a sandbox denial, whether it starts or reaches infrastructure or is a scenario's own command, re-run it via the Bash tool (`dangerouslyDisableSandbox: true`) before treating the infrastructure as unavailable or recording the scenario as failed, blocked, or inconclusive.

Isolation

  • Isolate shared process state so concurrent or subagent runs don't collide: bind dev servers and services to unique ports, scope tmux sessions (`tmux -L <name>`), give each browser session a unique name so cleanup can target only its own, and write screenshots and other scratch state to absolute paths under a uniquely named subdirectory of the session scratchpad directory. Derive each such identifier once and reuse that exact value in every later command, writing it as a literal or reading it back from a note in that subdirectory. A value recomputed per shell, such as `$$`, differs between the command that creates a resource and the command that releases it, so cleanup releases something it never created and reports success while the real resource leaks. A port picked as unique may already be held by a concurrent agent, so check it before binding and move to another when it is taken, leaving the incumbent running. When a unique port moves a service off its default address, find the settings elsewhere in the stack that name that default, such as allowed origins and sign-in callback URLs, and bring each in line through runtime overrides, leaving the working tree unchanged: add the new address beside the default in a list, and replace the default only where no process outside this run reads the setting.
  • Reuse a running dev server only when this session started it. Otherwise start one on a port this run selected and wait for it to be ready. Confirm it bound to that port before sending it traffic — a failed bind leaves another agent's service answering. Move to another port when the port is taken; report the error and stop when the server itself failed to start.
  • Run each test runner this run starts in its own process group under a timeout enforced from outside the runner.

Unavailable Infrastructure

When a scenario needs infrastructure that cannot be stood up in this session (backend service, auth provider, external dependency), look for a stub before treating the scenario as blocked. When the dependency is reached through a client whose endpoint is runtime configuration, repoint that endpoint at a local stub; the same interception yields an artifact the system emits rather than provides (a token, a session identifier, a single-use link). Confine this to runtime configuration and leave the working tree unchanged. When the code under test makes the call itself, a **call-site stub** intercepts it without a second process: from the run, install a replacement into the running process that matches one destination and returns the response the scenario needs or delays the real one, gated on a flag the run can flip. Installing it from the run leaves the working tree unchanged as well. It lasts only as long as the process, so re-install it after anything that restarts or reloads it. When no stub applies or the ones that do fail, the scenario is blocked: name what is missing and what was tried.

Privileged State and Second Participants

When a scenario needs privileged state or a second participant (an entitlement or plan tier, an elevated role, seed data, a second concurrent client or session), provision it through a path the project already exposes for development, such as its own development-only endpoint, an administrative command, or a second client this run starts. When the provisioning path writes to a shared external system, carry it through the write sequence under Writes to Shared External Systems. Treat the scenario as blocked only after an attempt to provision failed, naming the precondition and what was tried.

Treat a permission or scope granted mid-run to unblock a scenario as something this run created: before reporting cleanup complete, verify the production code never needs it, then ask for it to be revoked in the report. Name the call sites checked there too.

Writes to Shared External Systems

When a scenario writes to a shared external system and those writes are not cleanly undoable, run it against fixture data with nothing to act on whenever its expected outcome can still be observed that way, so the run still exercises wiring, auth, queries, guards, and failure isolation while writing nothing, and name that fixture data in the result as a substitution. Treat writes as not cleanly undoable whenever restoring the records leaves downstream effects the writes triggered in place.

When the writing path must run, work through it in order. Determine the full write set without executing it: use a dry-run mode when one exists, otherwise trace the code path and enumerate every record it writes, including those reached through triggers, cascades, and hooks. State what the enumeration cannot settle rather than presenting it as complete. Pick the target whose writes are incidental to what the scenario verifies, weighing each candidate's write set against the coverage it adds. Then request approval via `AskUserQuestion`, presenting the enumeration as what is being consented to, and request it again whenever the enumeration changes. Capture a pre-run ma

Read more
Ships withturbo

Reusable workflows for planning, building, reviewing, and shipping with Claude Code and Codex.

Get the whole plugin
Stats
405
Stars
32
Forks
Active
Maintenance
Python
Language
MIT
License
14h ago
Last commit
6mo ago
Created

Repo: tobihagemann/turbo

Other skills on turbo.