academic-slides
Use this skill for creating or refining an academic slide deck and the talk built around it:…
Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research. Apply it to
$ npx -y skills add evoscientist/evoskills --skill experiment-craft --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/experiment-craftContext preview
The summary Claude sees to decide when to auto-load this skill.
Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research. Apply it to
name: experiment-craft description: "Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research. Apply it to underperforming methods, training that will not converge, regressions after a change, inconsistent results across datasets, aimless experimentation without progress, and questions like 'why doesn't this work?', 'no progress after many attempts', or 'how should I investigate this failure?'. Also use it for setting up practical experiment logging/record-keeping that supports debugging and iteration. Do not use it for designing a brand-new experiment pipeline or full experiment program (use experiment-pipeline), generating research ideas, fixing isolated coding/syntax errors, or writing retrospective summaries into research memory/notes/knowledge bases." allowed-tools: "write_file edit_file read_file think_tool execute" metadata: author: EvoScientist version: '1.0.0' tags: [core, experimentation, experiment-design]
A systematic approach to running, debugging, and iterating on research experiments. The critical skill is not running more experiments — it's understanding WHY experiments fail.
> This skill is typically loaded from within `experiment-pipeline` when a stage attempt fails. After debugging, return to the pipeline's stage-gate structure to continue. Can also be used standalone for any experiment debugging.
**Finding WHY experiments fail is the most critical research skill.** Not analyzing results leads to two failure modes: 1. **Slow progress**: Running random experiments without understanding failure causes 2. **Wasted time**: Abandoning good approaches because activation tricks were missed
The goal is not to run more experiments. The goal is to run the RIGHT experiments — ones that isolate causes and test specific hypotheses.
When an experiment fails or produces unexpected results, follow these five steps:
Gather concrete examples of bad results. Look at the actual outputs, not just aggregate metrics. What specifically went wrong? Are the failures systematic or random?
You need a baseline that works. Two ways to find one:
If you can't find any working version, simplify further until something works. There is always a simple enough version that works.
Starting from the working version, incrementally add complexity until it breaks:
This step isolates the cause. Without it, you're guessing.
Based on the isolated cause from Step 3:
1. List possible explanations for why this factor causes failure 2. Rank by likelihood (based on your understanding and literature) 3. Design targeted experiments to verify or eliminate each hypothesis 4. Confirm the actual cause experimentally — don't rely on intuition alone
Based on the confirmed cause:
See [references/debugging-methodology.md](references/debugging-methodology.md) for detailed branching logic and a cause taxonomy.
Prioritize these rules during experimental work:
1. **Change only one variable at a time**: If you change two things and it works, you don't know which one fixed it. If you change two things and it doesn't work, you don't know which one is wrong. Single-variable changes are slower per experiment but faster overall. 2. **Fast iteration requires effective experiments, not more experiments**: Blind experimentation makes things worse. One well-designed diagnostic experiment is worth ten random trials. 3. **Some great techniques don't work alone**: They need specific activation tricks — learning rate schedules, initialization schemes, data preprocessing steps. Don't discard a technique after one failed attempt. Check related papers for their undisclosed tricks. 4. **Check related papers for their tricks**: Papers solving similar technical challenges often have critical implementation details buried in supplementary material or code. These tricks can make the difference between a technique working or failing. 5. **"Once you've ruled out the impossible, whatever remains must be true"**: Systematic elimination beats intuition. When debugging, explicitly list ALL possible causes, then eliminate them one by one with targeted experiments.
Every experiment should be logged with five sections. Use the template at [assets/experiment-log-template.md](assets/experiment-log-template.md).
| Section | What to Record | |---------|---------------| | Purpose | Why you're running this experiment; what y
The official skill repository for EvoScientist. Each skill is an installable knowledge pack that extends EvoScientist with domain-specific expertise.
Use this skill for creating or refining an academic slide deck and the talk built around it:…
Manages persistent research memory across ideation and experimentation cycles. Maintains two…
Use this skill whenever the user submits a non-trivial mathematical claim that needs a…
Iterative code refinement through plan → code → evaluate → refine cycles. Runs lint checks…
Guides structured 4-stage experiment execution with attempt budgets and gate conditions:…
Generate professional presentation slides and high-quality illustrations using Gemini image…