/benchmark
Use when timing a program or command, asking whether a change or version made it faster, or building a benchmark harness. For repeated optimization toward a target, use performance:hill-climb.
$ npx -y skills add bendrucker/claude --skill benchmark --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/benchmark
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when timing a program or command, asking whether a change or version made it faster, or building a benchmark harness. For repeated optimization toward a target, use performance:hill-climb.
SKILL.md
benchmark.SKILL.mdname: performance:benchmark
description: >-
Use when timing a program or command, asking whether a change or version
made it faster, or building a benchmark harness. For repeated optimization
toward a target, use performance:hill-climb.
argument-hint: "[<command>]"
Benchmark
Goal: a comparison the user can trust. Every arm runs the same scenario from the same starting state, the arms are interleaved so machine drift hits them equally, the noise floor is measured before any difference is read, and a difference counts only when it clears that floor.
Tools
- `hyperfine`: !`hyperfine --version 2>/dev/null || echo "not found. Required unless the project's own tooling covers its job."`
- `bun` (runs `compare.ts`): !`bun --version 2>/dev/null || echo "not found. Required for compare.ts."`
Existing Tooling
Look for the project's own benchmarks before building one: package scripts, a `bench/` directory, benchmark tests, CI jobs that time something. Check each against the requirements the setup below is built to meet:
- **Scenario.** It runs the scenario the user cares about, from the same starting state every run.
- **Isolation.** Network, remote state, and caches are stubbed, held fixed, or reset per run.
- **Interleaving.** Arms alternate within one run or one process, so drift hits every arm.
- **Samples.** It keeps every run's measurement, enough of them to show an A/A noise floor, not only a mean.
- **Significance.** It reports a significance test, or its samples can feed `compare.ts report`.
Use the existing tooling when it meets every requirement. When it misses some, compose around it and keep its scenario: run its command as a `compare.ts` arm for interleaving and significance, or add a `--prepare` reset. Tell the user which requirement it missed and what you added.
Scenario
Pin the scenario before timing it. A run that changes its own input measures a different scenario each time.
- Reset state per run with `--prepare` (restore a fixture, clear or warm a cache), and one-time setup with `--setup`. Match the cache state to the scenario the user cares about. A cold-start metric needs a cold cache on every run.
- Stub or hold fixed everything outside the program: network, remote state, other processes' files. When the program syncs with a remote, point it at a local copy that stays unchanged.
- Pass `-N` when the command is a single executable with arguments, which removes the shell's startup from every sample.
- Before a baseline, put the machine on AC power with power saving off, and quiet it as far as it goes. When it cannot go idle, interleaving keeps the comparison fair, and the load average `compare.ts run` prints each round explains a round that stands out.
- When the program is a test suite, raise the runner's timeout in every arm's command (`bun test --timeout`, `pytest --timeout`), so a load spike slows a run instead of failing it. Read the first timeout before raising it, since it can be a real cost worth a candidate.
- For a shell script or shell startup (`.zshrc`, `.bashrc`), read [references/shell.md](references/shell.md).
Arms
Interleaving runs every arm in the same round, so each one exists on disk at once:
- **Copies.** A `git worktree add <dir> <ref>` per version, or a separate build output per arm. Put each worktree outside the repo, where a test runner, linter, or watcher walking the tree cannot pick it up.
- **State.** When the program writes outside its tree (`~/.cache`, a config directory), point each arm at its own copy. Shared state lets a candidate that changes a cache format make both arms miss on every run.
- **Paths.** Check that each arm runs only its own copy. A test runner reads a bare path argument as a filter that can match another arm's tree, so pass paths with a leading `./`.
- **Passing run.** Before a long comparison, run each arm once with its output visible and confirm it passes.
Compare with the bundled script, which runs short `hyperfine` rounds in rotating order and pools them:
bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts run tmp/bench/<comparison> \
--arm 'base=<command>' --arm 'candidate=<command>' --rounds 6 --runs 3 -- -N --prepare '<reset>'
bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts report tmp/bench/<comparison> --markdownThe first `--arm` is the base. Arguments after `--` go to every `hyperfine` call. Give each comparison its own directory, since `run` refuses one that already holds exports. `report` also reads `--export-json` files from a plain `hyperfine` run, and pools every file and directory it is given per arm name.
Noise Floor
Before comparing versions, run the unchanged program as two arms (an A/A comparison) and repeat until the pair ties. When it stars, the harness is noisier than the effect:
- Raise `--runs` or `--rounds`.
- Quiet the machine.
- Find the state that leaks between runs.
The A/A spread (`±MAD`) is the smallest change the harness can resolve.
When the Mac's A/A will not tie or the fast signal needs `perf` or hardware counters, read [../profile/references/linux-vm.md](../profile/references/linux-vm.md).
Reading the Report
- `*` marks a change with permutation p below `--alpha` (0.1) and a size of at least `--min-effect` (3%). An unstarred change is a tie, however large it looks.
- `†` marks fewer than 4 runs on a side, too few to star. With 3 a side, the permutation test cannot reach p < 0.1 however far apart the arms are, so plan `--rounds` and `--runs` for at least 4 per arm, even when each run is expensive.
- A nonzero `failed` count means some runs exited nonzero and left the pool. `compare.ts run` stops after the round where a run fails. Run that arm once with its output visible and fix the cause before starting a new comparison.
- About one comparison in ten stars by chance at the default alpha. Re-run a star on a metric nobody predicted would move before treating it as a result.
Fast Signals
Wall time is the metric users feel, and the noi
Read more
name: performance:benchmark description: >- Use when timing a program or command, asking whether a change or version made it faster, or building a benchmark harness. For repeated optimization toward a target, use performance:hill-climb. argument-hint: "[<command>]"
Benchmark
Goal: a comparison the user can trust. Every arm runs the same scenario from the same starting state, the arms are interleaved so machine drift hits them equally, the noise floor is measured before any difference is read, and a difference counts only when it clears that floor.
Tools
- `hyperfine`: !`hyperfine --version 2>/dev/null || echo "not found. Required unless the project's own tooling covers its job."`
- `bun` (runs `compare.ts`): !`bun --version 2>/dev/null || echo "not found. Required for compare.ts."`
Existing Tooling
Look for the project's own benchmarks before building one: package scripts, a `bench/` directory, benchmark tests, CI jobs that time something. Check each against the requirements the setup below is built to meet:
- **Scenario.** It runs the scenario the user cares about, from the same starting state every run.
- **Isolation.** Network, remote state, and caches are stubbed, held fixed, or reset per run.
- **Interleaving.** Arms alternate within one run or one process, so drift hits every arm.
- **Samples.** It keeps every run's measurement, enough of them to show an A/A noise floor, not only a mean.
- **Significance.** It reports a significance test, or its samples can feed `compare.ts report`.
Use the existing tooling when it meets every requirement. When it misses some, compose around it and keep its scenario: run its command as a `compare.ts` arm for interleaving and significance, or add a `--prepare` reset. Tell the user which requirement it missed and what you added.
Scenario
Pin the scenario before timing it. A run that changes its own input measures a different scenario each time.
- Reset state per run with `--prepare` (restore a fixture, clear or warm a cache), and one-time setup with `--setup`. Match the cache state to the scenario the user cares about. A cold-start metric needs a cold cache on every run.
- Stub or hold fixed everything outside the program: network, remote state, other processes' files. When the program syncs with a remote, point it at a local copy that stays unchanged.
- Pass `-N` when the command is a single executable with arguments, which removes the shell's startup from every sample.
- Before a baseline, put the machine on AC power with power saving off, and quiet it as far as it goes. When it cannot go idle, interleaving keeps the comparison fair, and the load average `compare.ts run` prints each round explains a round that stands out.
- When the program is a test suite, raise the runner's timeout in every arm's command (`bun test --timeout`, `pytest --timeout`), so a load spike slows a run instead of failing it. Read the first timeout before raising it, since it can be a real cost worth a candidate.
- For a shell script or shell startup (`.zshrc`, `.bashrc`), read [references/shell.md](references/shell.md).
Arms
Interleaving runs every arm in the same round, so each one exists on disk at once:
- **Copies.** A `git worktree add <dir> <ref>` per version, or a separate build output per arm. Put each worktree outside the repo, where a test runner, linter, or watcher walking the tree cannot pick it up.
- **State.** When the program writes outside its tree (`~/.cache`, a config directory), point each arm at its own copy. Shared state lets a candidate that changes a cache format make both arms miss on every run.
- **Paths.** Check that each arm runs only its own copy. A test runner reads a bare path argument as a filter that can match another arm's tree, so pass paths with a leading `./`.
- **Passing run.** Before a long comparison, run each arm once with its output visible and confirm it passes.
Compare with the bundled script, which runs short `hyperfine` rounds in rotating order and pools them:
bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts run tmp/bench/<comparison> \
--arm 'base=<command>' --arm 'candidate=<command>' --rounds 6 --runs 3 -- -N --prepare '<reset>'
bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts report tmp/bench/<comparison> --markdownThe first `--arm` is the base. Arguments after `--` go to every `hyperfine` call. Give each comparison its own directory, since `run` refuses one that already holds exports. `report` also reads `--export-json` files from a plain `hyperfine` run, and pools every file and directory it is given per arm name.
Noise Floor
Before comparing versions, run the unchanged program as two arms (an A/A comparison) and repeat until the pair ties. When it stars, the harness is noisier than the effect:
- Raise `--runs` or `--rounds`.
- Quiet the machine.
- Find the state that leaks between runs.
The A/A spread (`±MAD`) is the smallest change the harness can resolve.
When the Mac's A/A will not tie or the fast signal needs `perf` or hardware counters, read [../profile/references/linux-vm.md](../profile/references/linux-vm.md).
Reading the Report
- `*` marks a change with permutation p below `--alpha` (0.1) and a size of at least `--min-effect` (3%). An unstarred change is a tie, however large it looks.
- `†` marks fewer than 4 runs on a side, too few to star. With 3 a side, the permutation test cannot reach p < 0.1 however far apart the arms are, so plan `--rounds` and `--runs` for at least 4 per arm, even when each run is expensive.
- A nonzero `failed` count means some runs exited nonzero and left the pool. `compare.ts run` stops after the round where a run fails. Run that arm once with its output visible and fix the cause before starting a new comparison.
- About one comparison in ten stars by chance at the default alpha. Re-run a star on a metric nobody predicted would move before treating it as a result.
Fast Signals
Wall time is the metric users feel, and the noi
My personal plugin marketplace for Claude Code, Anthropic's AI coding assistant.
Repo: bendrucker/claude
Other skills on bendrucker-claude.
cli
Type-safe CLI argument parsing with commander and @commander-js/extra-typings, the standard…
activity
Report real device usage from ActivityWatch. Covers per-app time, window titles, and active…
history
Report shell history from atuin's local capture. Covers what commands ran, when, where, and…
bun
Bun runtime patterns. Use when running bun commands, working with package.json/bun.lock,…
agent-team
Orchestrating Claude Code agent teams. Use when creating teams, spawning teammates, assigning…

