Skip to content
Automation
Skill

/experiment-suite

Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.

From plugin
ai4s-skills
2257 skills
Install
$ npx -y skills add ai4s-research/ai4s-skills --skill experiment-suite --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/experiment-suite

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.

SKILL.md

experiment-suite.SKILL.md
name: experiment-suite
description: Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.

Experiment Suite

Overview

End-to-end experiment package builder. **Single stage, full quality from the start.** The agent (Claude Code / Cursor / Aider / Codex / …) writes everything directly using its own tools (Write, Bash, WebFetch, …). This skill contains procedure + reference playbooks + figure-example scripts — no Python runtime, no LLM SDK.

The substantive work is decomposed into reference playbooks under `references/`:

| Reference | Topic | |---|---| | `references/00-incremental-execution.md` | how to do this without losing work: batches, persistence, resume — **read first** | | `references/01-design-depth.md` | what a real experiment design contains (motivation → hypothesis → datasets → baselines → metrics → ablations → budget) | | `references/01a-data-contract.md` | runtime dataset binding: source, access route, version, split, and reuse boundary | | `references/02-code-quality.md` | code-skeleton standards — runnable `model.py`, `data.py`, `train.py`, `evaluate.py` | | `references/03-results-protocol.md` | `results.json` schema; `measured` / `simulated` / `illustrative` provenance | | `references/04-publication-figures.md` | publication-grade charts, multi-panel layouts, taste rules | | `references/04a-figure-contract.md` | figure logic before plotting: conclusion, panel map, reviewer risk | | `references/04b-figure-qa.md` | export bundle, editable text, statistics and image-integrity QA | | `references/05-report-structure.md` | structured `experiment_report.md` (problem → design → method → results → analysis → limitations) | | `references/06-quality-gate.md` | self-check before delivery |

Also: `figure_examples/` — publication-style matplotlib scripts plus a shared style kit the agent can use as starting points.

**Read the relevant reference _before_ writing, not after.** The full pass does not fit in a single turn — `references/00-incremental-execution.md` is the only execution mode that completes.

When to Use

  • User wants to "design an experiment" for a research question.
  • User needs runnable code for a specific task (classification / forecasting / detection / …).
  • User wants to compare methods and have a structured report at the end.
  • User needs publication-quality figures of experimental results.

When NOT to Use

  • User only wants a quick code snippet (write code directly).
  • User wants a full paper → `paper-writer`.
  • User wants a literature survey → `literature-survey`.

Workflow

Step 1 — Understand the question and operating mode

Confirm with the user:

  • **Research question** — what we are trying to answer.
  • **Task type** — classification / regression / forecasting / detection / generation / …
  • **Mode**
  • **measured** — user has real data or will run code themselves; provide a path to a measured `results.json` or run `train.py` against real data later.
  • **simulated** (default) — agent generates a plausible-shaped, deterministic `results.json` as a placeholder. Every figure/table caption must say "simulated".
  • **Framework preference** — PyTorch (default), JAX, TensorFlow, or sklearn.
  • **Compute budget** — hours / GPUs available; constrains the code skeleton and hyperparameter plan.

If the user has data and time, push toward measured mode. If not, simulated is acceptable **provided** disclosures are honest in every artefact.

Step 2 — Set up the run directory

QUESTION="<research_question>"
SLUG=$(python3 -c "import re,hashlib,sys; t=sys.argv[1]; n=re.sub(r'[\\s_]+','-',re.sub(r'[^\\w\\s-]','',t.lower().strip())).strip('-')[:40].rstrip('-'); h=hashlib.sha1(t.encode()).hexdigest()[:8]; print(f'{n}-{h}')" "$QUESTION")
TS=$(date +%Y-%m-%d_%H%M%S)
RUN=output/experiment-suite/$SLUG/$TS

mkdir -p "$RUN/experiment" "$RUN/figures"
ln -sfn "$TS" "output/experiment-suite/$SLUG/latest"

In commands below `$RUN` = `output/experiment-suite/<slug>/latest`.

The agent will create five top-level files inside `$RUN/`:

  • `experiment_design.md`
  • `data_contract.md`
  • `experiment/{model.py,data.py,train.py,evaluate.py,config.yaml,requirements.txt,README.md}`
  • `results.json`
  • `figures/*.pdf` plus their `make_*.py` source and a `manifest.json`
  • `experiment_report.md`

Step 3 — Build the package (REQUIRED — this is the whole job)

Open `references/00-incremental-execution.md` first. Then carry out the six tracks below across many turns, persisting state to `$RUN/` after every batch.

3.1 Design — full justification document

**Open:** `references/01-design-depth.md` and `references/01a-data-contract.md`. First write `$RUN/data_contract.md` as the dataset contract for this run. It must say whether the data are user-supplied, agent-discovered, reused public, controlled, or synthetic fallback. Then write `$RUN/experiment_design.md` as a real design (≥ 700 words): motivation → hypothesis → datasets → baselines → metrics → ablations → compute budget. Justify every choice.

3.2 Code — actually runnable

**Open:** `references/02-code-quality.md`. Fill `$RUN/experiment/` with code that an engineer could launch with `python train.py --config config.yaml`. Real (if minimal) model class, real data loader, real train loop, real eval. The generated `data.py` and `config.yaml` are runtime products of this run and should bind to `$RUN/data_contract.md`, not to a repository-wide hard-coded benchmark. Add a `README.md` with run instructions.

3.3 Results — honest provenance

**Open:** `references/03-results-protocol.md`. Produce `$RUN/results.json` with a well-formed schema: per-seed entries, per-method per-metric mean & std, ablation block, and a `provenance` field that names the source.

  • **measured mode** — the user runs `experiment/tr
Read more
Ships withai4s-skills

Open-source agent skills for AI for Science: topic exploration, literature survey, experiments, paper writing, and integrity audit — driven by any coding agent.

Get the whole plugin
Stats
225
Stars
21
Forks
Maintained
Maintenance
Python
Language
MIT
License
1mo ago
Last commit
2mo ago
Created

Repo: ai4s-research/ai4s-skills

Other skills on ai4s-skills.