Skip to content
Agent Orchestration
Skill

/loopx-benchmark

Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task

BOOST
From plugin
loopx
6.2k13 skills1 command
Install
$ npx -y skills add loopx-project/loopx --skill loopx-benchmark --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/loopx-benchmark

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task

SKILL.md

loopx-benchmark.SKILL.md
name: loopx-benchmark
description: "Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task under an existing runner does not by itself select this skill. Excludes casual benchmark discussion and ordinary software microbenchmarks."

LoopX Benchmark Workflow

Use this skill to operate or analyze a LoopX-managed benchmark experiment. The builtin `benchmark-toolkit` capability owns provider-neutral experiment state and integrity boundaries. This packaged skill is its task-triggered Agent playbook.

An assigned solver follows its task instructions and current execution contract. The words “benchmark”, “evaluation”, or “submission” in that task do not grant the operator role or require experiment-board discovery. Use this workflow when the requested work actually includes run management or post-run analysis; an explicit request to use the skill still applies. Keep the solver's task-local validation and authorized submission path distinct from experiment management.

The capability is catalog-ready without a per-Goal enable switch. Installing this skill does not grant runner, shell, network, credential, private-evidence, or Goal mutation authority. Respect the selected todo's required capabilities, any external provider binding, host permissions, and user gates.

Capability surface

  • `loopx capability show benchmark-toolkit --format json` — catalog entry with

usage hints, role boundaries, and the post-run case-insight template.

  • `loopx benchmark --help` — subcommands (experiment-board-show,

experiment-board-upsert, source-revision-fence, integrity-qualification, classify-artifacts).

Share a study through the public-safe contract

When another benchmark developer needs portable study data, use the capability's typed study flow rather than sharing a runner-specific ledger or raw evidence:

1. Validate `benchmark_study_manifest_v0` with `benchmark study-validate`. 2. Wrap one allowlisted manifest, experiment-board row, redacted insight, or runtime observation with `benchmark upload-envelope`. 3. Run `benchmark upload-local` without `--execute` first, then explicitly execute against a caller-selected local JSONL store. 4. Verify the record/digest/revision binding with `benchmark upload-readback`. 5. Derive the campaign/arm/case/run packet with `benchmark study-dashboard`; pass a compact four-arm contract only when the study preregistered that design.

For `case_insight_projection`, first upload the same run's active terminal experiment-board row with `insight.status=complete`. The case, run, and outcome must match; the run row remains the only arm, score, countability, integrity, and treatment-fidelity authority. Reduce private post-run evidence to bounded prose and public-safe handles or digests before building the envelope.

The local provider is a no-network simulation. It does not grant remote upload, publication, credentials, retention, or benchmark submission authority. Adapters keep their native metric names and reduce private post-run evidence before envelope construction.

Share exploratory behavior findings

When the owner authorizes selected behavioral observations but not a complete study release, use `behavior_finding` records and `benchmark behavior-report`. See `docs/reference/benchmark-behavior-findings.md` for the contract. These records require selection rules, sample denominators, observations, interpretations, limitations, counterevidence, and evidence digests; they require neither a run-row upload nor a full study manifest and have no score authority.

Freeze the authorized disclosure projection before rendering. Review the same scope in visible text, foldouts, embedded data, downloads, and PR attachments. Permission to share duration does not grant permission to share outcome totals or deltas. Schema validity and a producer redaction attestation are not publication approval or verification of unshared evidence. Keep selected-case observations explicitly exploratory and retain the relevant limitations and counterexamples.

Select the operating lane

  • **Inspect or explain:** use `capability show` and `benchmark --help`; remain

read-only. Do not create an experiment-board row merely because the user asks what the toolkit does.

  • **Plan, select, or launch a run:** follow the experiment sequence below. The

first action is to read the board; a launch still requires an authorized runner and admitted source.

  • **Monitor an active campaign:** read the board and runtime-owned projections;

update only on material run transitions. Do not manufacture progress from a timer tick.

  • **Analyze a terminal run:** wait until solving is terminal and scoring is

complete before reading hidden evaluator evidence or writing a case insight.

For a generic library microbenchmark or an eval with no LoopX Goal/board, use the task's normal tools instead of imposing this workflow.

Experiment sequence

1. **Read the experiment board before launching or selecting a case.**

   loopx benchmark experiment-board-show --goal-id <GOAL_ID> --format json

Inspect baseline, treatment, explore, countability, effort, and insight rows before choosing the next arm.

2. **Qualify the source revision before each new run admission.**

   loopx benchmark source-revision-fence \
     --source-checkout <clean-source> \
     --expected-revision <PIN> \
     --observed-reference-revision <OBSERVED_HEAD> \
     --require-admitted --format json

The fence fails closed unless the clean pinned source matches the observed reference head.

3. **Preview, then preregister or mark the run row when it starts.**

   loopx benchmark experiment-board-upsert --goal-id <GOAL_ID> \
     --row-json <running-row.json
Read more
Ships withloopx

A control plane with a durable state kernel for long-horizon agents and teams. Keep work moving and improving across sessions, with less human attention.

Get the whole plugin
Stats
6,202
Stars
596
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
10m ago
Last commit
4mo ago
Created
23h ago
Added

Repo: loopx-project/loopx

Other skills on loopx.