swe-marathon-five-arm
Migrate historical SWE-Marathon agent configurations to the shared Codex runtime.
Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task
$ npx -y skills add loopx-project/loopx --skill loopx-benchmark --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/loopx-benchmarkContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task
name: loopx-benchmark description: "Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment through benchmark-toolkit: select or launch runs, maintain experiment-board rows, qualify integrity, or analyze matched comparisons and case insights. Solving an assigned benchmark task under an existing runner does not by itself select this skill. Excludes casual benchmark discussion and ordinary software microbenchmarks."
Use this skill to operate or analyze a LoopX-managed benchmark experiment. The builtin `benchmark-toolkit` capability owns provider-neutral experiment state and integrity boundaries. This packaged skill is its task-triggered Agent playbook.
An assigned solver follows its task instructions and current execution contract. The words “benchmark”, “evaluation”, or “submission” in that task do not grant the operator role or require experiment-board discovery. Use this workflow when the requested work actually includes run management or post-run analysis; an explicit request to use the skill still applies. Keep the solver's task-local validation and authorized submission path distinct from experiment management.
The capability is catalog-ready without a per-Goal enable switch. Installing this skill does not grant runner, shell, network, credential, private-evidence, or Goal mutation authority. Respect the selected todo's required capabilities, any external provider binding, host permissions, and user gates.
usage hints, role boundaries, and the post-run case-insight template.
experiment-board-upsert, source-revision-fence, integrity-qualification, classify-artifacts).
When another benchmark developer needs portable study data, use the capability's typed study flow rather than sharing a runner-specific ledger or raw evidence:
1. Validate `benchmark_study_manifest_v0` with `benchmark study-validate`. 2. Wrap one allowlisted manifest, experiment-board row, redacted insight, or runtime observation with `benchmark upload-envelope`. 3. Run `benchmark upload-local` without `--execute` first, then explicitly execute against a caller-selected local JSONL store. 4. Verify the record/digest/revision binding with `benchmark upload-readback`. 5. Derive the campaign/arm/case/run packet with `benchmark study-dashboard`; pass a compact four-arm contract only when the study preregistered that design.
For `case_insight_projection`, first upload the same run's active terminal experiment-board row with `insight.status=complete`. The case, run, and outcome must match; the run row remains the only arm, score, countability, integrity, and treatment-fidelity authority. Reduce private post-run evidence to bounded prose and public-safe handles or digests before building the envelope.
The local provider is a no-network simulation. It does not grant remote upload, publication, credentials, retention, or benchmark submission authority. Adapters keep their native metric names and reduce private post-run evidence before envelope construction.
When the owner authorizes selected behavioral observations but not a complete study release, use `behavior_finding` records and `benchmark behavior-report`. See `docs/reference/benchmark-behavior-findings.md` for the contract. These records require selection rules, sample denominators, observations, interpretations, limitations, counterevidence, and evidence digests; they require neither a run-row upload nor a full study manifest and have no score authority.
Freeze the authorized disclosure projection before rendering. Review the same scope in visible text, foldouts, embedded data, downloads, and PR attachments. Permission to share duration does not grant permission to share outcome totals or deltas. Schema validity and a producer redaction attestation are not publication approval or verification of unshared evidence. Keep selected-case observations explicitly exploratory and retain the relevant limitations and counterexamples.
read-only. Do not create an experiment-board row merely because the user asks what the toolkit does.
first action is to read the board; a launch still requires an authorized runner and admitted source.
update only on material run transitions. Do not manufacture progress from a timer tick.
complete before reading hidden evaluator evidence or writing a case insight.
For a generic library microbenchmark or an eval with no LoopX Goal/board, use the task's normal tools instead of imposing this workflow.
1. **Read the experiment board before launching or selecting a case.**
loopx benchmark experiment-board-show --goal-id <GOAL_ID> --format json
Inspect baseline, treatment, explore, countability, effort, and insight rows before choosing the next arm.
2. **Qualify the source revision before each new run admission.**
loopx benchmark source-revision-fence \
--source-checkout <clean-source> \
--expected-revision <PIN> \
--observed-reference-revision <OBSERVED_HEAD> \
--require-admitted --format jsonThe fence fails closed unless the clean pinned source matches the observed reference head.
3. **Preview, then preregister or mark the run row when it starts.**
loopx benchmark experiment-board-upsert --goal-id <GOAL_ID> \
--row-json <running-row.jsonA control plane with a durable state kernel for long-horizon agents and teams. Keep work moving and improving across sessions, with less human attention.
Repo: loopx-project/loopx
Migrate historical SWE-Marathon agent configurations to the shared Codex runtime.
在 Terminal-Bench 4.0 上做 codex harness 五臂对照(裸 codex / 原生 /goal / LoopX 三模式)。复用 SWE-Marathon…
Inspect authorized LoopX Goals, Todos and deliveries to explain progress, identify owner…
Qualify the exact final diff for a LoopX-managed goal. Use when goal policy enables…
Use when a connected LoopX project is asked to read, remember, record, index, register, or…
Operate an explicitly activated LoopX Material Lifecycle for a connected project. Use for…