ablation-planner
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.
$ npx -y skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill monitor-experiment --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/monitor-experimentContext preview
The summary Claude sees to decide when to auto-load this skill.
Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.
name: monitor-experiment description: Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output. argument-hint: "[server-alias or screen-name]" allowed-tools: Bash(ssh *), Bash(echo *), Read, Write, Edit
> ⏱ **External cadence is appropriate here.** This skill waits on an external > fact (job completion / progress), so it is a natural `/loop` / `CronCreate` > surface: the wake reads status and self-judges only **machine-checkable** > completion (exit code, file exists, epoch logged) — never quality. This is > the additive external-wait shape in > [`shared-references/external-cadence.md`](../shared-references/external-cadence.md). > If a scheduled wait here ends in a verdict step (e.g. then audit results), > run that verdict **once** after the wait clears — not re-entered per tick.
Monitor: $ARGUMENTS
**SSH server:**
ssh <server> "screen -ls"
**Vast.ai instance** (read `ssh_host`, `ssh_port` from `vast-instances.json`):
ssh -p <PORT> root@<HOST> "screen -ls"
Also check vast.ai instance status:
vastai show instances
**Modal** (when `gpu: modal` in CLAUDE.md):
modal app list # List running/recent apps modal app logs <app> # Stream logs from a running app
Modal apps auto-terminate when done — if it's not in the list, it already finished. Check results via `modal volume ls <volume>` or local output.
For each screen session, capture the last N lines:
ssh <server> "screen -S <name> -X hardcopy /tmp/screen_<name>.txt && tail -50 /tmp/screen_<name>.txt"
If hardcopy fails, check for log files or tee output.
ssh <server> "ls -lt <results_dir>/*.json 2>/dev/null | head -20"
If JSON results exist, fetch and parse them:
ssh <server> "cat <results_dir>/<latest>.json"
**Skip this step entirely if `wandb` is not set or is `false` in CLAUDE.md.**
Pull training curves and metrics from Weights & Biases via Python API:
# List recent runs in the project
ssh <server> "python3 -c \"
import wandb
api = wandb.Api()
runs = api.runs('<entity>/<project>', per_page=10)
for r in runs:
print(f'{r.id} {r.state} {r.name} {r.summary.get(\"eval/loss\", \"N/A\")}')
\""
# Pull specific metrics from a run (last 50 steps)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
history = list(run.scan_history(keys=['train/loss', 'eval/loss', 'eval/ppl', 'train/lr'], page_size=50))
print(json.dumps(history[-10:], indent=2))
\""
# Pull run summary (final metrics)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
print(json.dumps(dict(run.summary), indent=2, default=str))
\""**What to extract:**
**W&B dashboard link** (include in summary for user):
https://wandb.ai/<entity>/<project>/runs/<run_id>
> This gives the auto-review-loop richer signal than just screen output — training dynamics, loss curves, and metric trends over time.
Present results in a comparison table:
| Experiment | Metric | Delta vs Baseline | Status | |-----------|--------|-------------------|--------| | Baseline | X.XX | — | done | | Method A | X.XX | +Y.Y | done |
After results are collected, check `~/.claude/feishu.json`:
· · · · · · -orange?style=flat) · · 💬 Join Community · 💡 Use ARIS as a skill-based workflow in Claude Code / Codex CLI / Cursor / Trae / Antigravity / GitHub Copilot CLI / OpenClaw / DeepSeek Harness, or get the full experience with the standalone ARIS-Code
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Quick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes…
Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to…
Search, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper…
Autonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\",…
Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop…