Skip to content
Development
Skill

/benchmark

Use when timing a program or command, asking whether a change or version made it faster, or building a benchmark harness. For repeated optimization toward a target, use performance:hill-climb.

BOOST
From plugin
bendrucker-claude
1788 skills10 agents1 MCP
Install
$ npx -y skills add bendrucker/claude --skill benchmark --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/benchmark

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when timing a program or command, asking whether a change or version made it faster, or building a benchmark harness. For repeated optimization toward a target, use performance:hill-climb.

SKILL.md

benchmark.SKILL.md
name: performance:benchmark
description: >-
  Use when timing a program or command, asking whether a change or version
  made it faster, or building a benchmark harness. For repeated optimization
  toward a target, use performance:hill-climb.
argument-hint: "[<command>]"

Benchmark

Goal: a comparison the user can trust. Every arm runs the same scenario from the same starting state, the arms are interleaved so machine drift hits them equally, the noise floor is measured before any difference is read, and a difference counts only when it clears that floor.

Tools

  • `hyperfine`: !`hyperfine --version 2>/dev/null || echo "not found. Required unless the project's own tooling covers its job."`
  • `bun` (runs `compare.ts`): !`bun --version 2>/dev/null || echo "not found. Required for compare.ts."`

Existing Tooling

Look for the project's own benchmarks before building one: package scripts, a `bench/` directory, benchmark tests, CI jobs that time something. Check each against the requirements the setup below is built to meet:

  • **Scenario.** It runs the scenario the user cares about, from the same starting state every run.
  • **Isolation.** Network, remote state, and caches are stubbed, held fixed, or reset per run.
  • **Interleaving.** Arms alternate within one run or one process, so drift hits every arm.
  • **Samples.** It keeps every run's measurement, enough of them to show an A/A noise floor, not only a mean.
  • **Significance.** It reports a significance test, or its samples can feed `compare.ts report`.

Use the existing tooling when it meets every requirement. When it misses some, compose around it and keep its scenario: run its command as a `compare.ts` arm for interleaving and significance, or add a `--prepare` reset. Tell the user which requirement it missed and what you added.

Scenario

Pin the scenario before timing it. A run that changes its own input measures a different scenario each time.

  • Reset state per run with `--prepare` (restore a fixture, clear or warm a cache), and one-time setup with `--setup`. Match the cache state to the scenario the user cares about. A cold-start metric needs a cold cache on every run.
  • Stub or hold fixed everything outside the program: network, remote state, other processes' files. When the program syncs with a remote, point it at a local copy that stays unchanged.
  • Pass `-N` when the command is a single executable with arguments, which removes the shell's startup from every sample.
  • Before a baseline, put the machine on AC power with power saving off, and quiet it as far as it goes. When it cannot go idle, interleaving keeps the comparison fair, and the load average `compare.ts run` prints each round explains a round that stands out.
  • When the program is a test suite, raise the runner's timeout in every arm's command (`bun test --timeout`, `pytest --timeout`), so a load spike slows a run instead of failing it. Read the first timeout before raising it, since it can be a real cost worth a candidate.
  • For a shell script or shell startup (`.zshrc`, `.bashrc`), read [references/shell.md](references/shell.md).

Arms

Interleaving runs every arm in the same round, so each one exists on disk at once:

  • **Copies.** A `git worktree add <dir> <ref>` per version, or a separate build output per arm. Put each worktree outside the repo, where a test runner, linter, or watcher walking the tree cannot pick it up.
  • **State.** When the program writes outside its tree (`~/.cache`, a config directory), point each arm at its own copy. Shared state lets a candidate that changes a cache format make both arms miss on every run.
  • **Paths.** Check that each arm runs only its own copy. A test runner reads a bare path argument as a filter that can match another arm's tree, so pass paths with a leading `./`.
  • **Passing run.** Before a long comparison, run each arm once with its output visible and confirm it passes.

Compare with the bundled script, which runs short `hyperfine` rounds in rotating order and pools them:

bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts run tmp/bench/<comparison> \
  --arm 'base=<command>' --arm 'candidate=<command>' --rounds 6 --runs 3 -- -N --prepare '<reset>'
bun ${CLAUDE_SKILL_DIR}/scripts/compare.ts report tmp/bench/<comparison> --markdown

The first `--arm` is the base. Arguments after `--` go to every `hyperfine` call. Give each comparison its own directory, since `run` refuses one that already holds exports. `report` also reads `--export-json` files from a plain `hyperfine` run, and pools every file and directory it is given per arm name.

Noise Floor

Before comparing versions, run the unchanged program as two arms (an A/A comparison) and repeat until the pair ties. When it stars, the harness is noisier than the effect:

  • Raise `--runs` or `--rounds`.
  • Quiet the machine.
  • Find the state that leaks between runs.

The A/A spread (`±MAD`) is the smallest change the harness can resolve.

When the Mac's A/A will not tie or the fast signal needs `perf` or hardware counters, read [../profile/references/linux-vm.md](../profile/references/linux-vm.md).

Reading the Report

  • `*` marks a change with permutation p below `--alpha` (0.1) and a size of at least `--min-effect` (3%). An unstarred change is a tie, however large it looks.
  • `†` marks fewer than 4 runs on a side, too few to star. With 3 a side, the permutation test cannot reach p < 0.1 however far apart the arms are, so plan `--rounds` and `--runs` for at least 4 per arm, even when each run is expensive.
  • A nonzero `failed` count means some runs exited nonzero and left the pool. `compare.ts run` stops after the round where a run fails. Run that arm once with its output visible and fix the cause before starting a new comparison.
  • About one comparison in ten stars by chance at the default alpha. Re-run a star on a metric nobody predicted would move before treating it as a result.

Fast Signals

Wall time is the metric users feel, and the noi

Read more
Ships withbendrucker-claude

My personal plugin marketplace for Claude Code, Anthropic's AI coding assistant.

Get the whole plugin

Other skills on bendrucker-claude.