ai4s-agent
Use when the user wants an end-to-end AI4S research pipeline — broad direction or specific topic in, full research package out (exploration + literature survey…
Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.
$ npx -y skills add ai4s-research/ai4s-skills --skill experiment-suite --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/experiment-suiteContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.
name: experiment-suite description: Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.
End-to-end experiment package builder. **Single stage, full quality from the start.** The agent (Claude Code / Cursor / Aider / Codex / …) writes everything directly using its own tools (Write, Bash, WebFetch, …). This skill contains procedure + reference playbooks + figure-example scripts — no Python runtime, no LLM SDK.
The substantive work is decomposed into reference playbooks under `references/`:
| Reference | Topic | |---|---| | `references/00-incremental-execution.md` | how to do this without losing work: batches, persistence, resume — **read first** | | `references/01-design-depth.md` | what a real experiment design contains (motivation → hypothesis → datasets → baselines → metrics → ablations → budget) | | `references/01a-data-contract.md` | runtime dataset binding: source, access route, version, split, and reuse boundary | | `references/02-code-quality.md` | code-skeleton standards — runnable `model.py`, `data.py`, `train.py`, `evaluate.py` | | `references/03-results-protocol.md` | `results.json` schema; `measured` / `simulated` / `illustrative` provenance | | `references/04-publication-figures.md` | publication-grade charts, multi-panel layouts, taste rules | | `references/04a-figure-contract.md` | figure logic before plotting: conclusion, panel map, reviewer risk | | `references/04b-figure-qa.md` | export bundle, editable text, statistics and image-integrity QA | | `references/05-report-structure.md` | structured `experiment_report.md` (problem → design → method → results → analysis → limitations) | | `references/06-quality-gate.md` | self-check before delivery |
Also: `figure_examples/` — publication-style matplotlib scripts plus a shared style kit the agent can use as starting points.
**Read the relevant reference _before_ writing, not after.** The full pass does not fit in a single turn — `references/00-incremental-execution.md` is the only execution mode that completes.
Confirm with the user:
If the user has data and time, push toward measured mode. If not, simulated is acceptable **provided** disclosures are honest in every artefact.
QUESTION="<research_question>"
SLUG=$(python3 -c "import re,hashlib,sys; t=sys.argv[1]; n=re.sub(r'[\\s_]+','-',re.sub(r'[^\\w\\s-]','',t.lower().strip())).strip('-')[:40].rstrip('-'); h=hashlib.sha1(t.encode()).hexdigest()[:8]; print(f'{n}-{h}')" "$QUESTION")
TS=$(date +%Y-%m-%d_%H%M%S)
RUN=output/experiment-suite/$SLUG/$TS
mkdir -p "$RUN/experiment" "$RUN/figures"
ln -sfn "$TS" "output/experiment-suite/$SLUG/latest"In commands below `$RUN` = `output/experiment-suite/<slug>/latest`.
The agent will create five top-level files inside `$RUN/`:
Open `references/00-incremental-execution.md` first. Then carry out the six tracks below across many turns, persisting state to `$RUN/` after every batch.
**Open:** `references/01-design-depth.md` and `references/01a-data-contract.md`. First write `$RUN/data_contract.md` as the dataset contract for this run. It must say whether the data are user-supplied, agent-discovered, reused public, controlled, or synthetic fallback. Then write `$RUN/experiment_design.md` as a real design (≥ 700 words): motivation → hypothesis → datasets → baselines → metrics → ablations → compute budget. Justify every choice.
**Open:** `references/02-code-quality.md`. Fill `$RUN/experiment/` with code that an engineer could launch with `python train.py --config config.yaml`. Real (if minimal) model class, real data loader, real train loop, real eval. The generated `data.py` and `config.yaml` are runtime products of this run and should bind to `$RUN/data_contract.md`, not to a repository-wide hard-coded benchmark. Add a `README.md` with run instructions.
**Open:** `references/03-results-protocol.md`. Produce `$RUN/results.json` with a well-formed schema: per-seed entries, per-method per-metric mean & std, ablation block, and a `provenance` field that names the source.
Open-source agent skills for AI for Science: topic exploration, literature survey, experiments, paper writing, and integrity audit — driven by any coding agent.
Use when the user wants an end-to-end AI4S research pipeline — broad direction or specific topic in, full research package out (exploration + literature survey…
Use when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report. Works…
Use when the user wants a comprehensive literature survey on a specific research topic. Outputs a complete PDF survey (6–20 pages, 60+ real citations, 100+…
Generate beautiful, high-resolution mindmaps from Markdown unordered lists. Outputs interactive HTML, HD PNG, and PDF with colorful branch themes.
Use when the user wants a complete, publication-grade research paper on a specific topic — produces 200+ real citations, 4–8 publication-grade figures, and 7…
Use when the user has a vague research direction and wants to explore feasible specific topics. Outputs a structured analysis with candidate topics,…