formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced…
Design fair comparison experiments against baselines and competing methods
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill comparison-design --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/comparison-designContext preview
The summary Claude sees to decide when to auto-load this skill.
Design fair comparison experiments against baselines and competing methods
name: comparison-design description: Design fair comparison experiments against baselines and competing methods version: 1.0.0 category: experiment-execution type: strategy sops: - baseline-selection - metric-specification - sample-size-estimation - seed-protocol-design - environment-specification tactics: - statistical-method-selection - reproducibility-protocol dependencies: sops: - baseline-selection - environment-specification - metric-specification - sample-size-estimation - seed-protocol-design tactics: - reproducibility-protocol - statistical-method-selection
**Question**: How much better is our method than the baseline?
1. **baseline-selection** → Select appropriate baselines (SOTA, simple, oracle) 2. **metric-specification** → Define primary metric and secondary metrics 3. **sample-size-estimation** → Power analysis for detecting meaningful differences 4. **seed-protocol-design** → Ensure fair random initialization across methods 5. **environment-specification** → Lock environment to prevent confounds 6. **reproducibility-protocol** (tactic) → Ensure all results are reproducible 7. **statistical-method-selection** (tactic) → Choose Bayesian or frequentist comparison
| Comparison Scope | Baselines | Datasets | Seeds | Min Runs | |-----------------|-----------|----------|-------|----------| | Minimal | 1 SOTA + 1 simple | 1 | 3 | 6 | | Standard | 2-3 baselines | 2-3 | 5 | 30-45 | | Comprehensive | 4+ baselines | 3-5 | 5-10 | 100+ | | Publication-ready | All relevant | 5+ | 10+ | 200+ |
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | reproducibility-protocol | Ensure experiment reproducibility through systematic environment and seed control | | statistical-method-selection | Select appropriate statistical methods for experiment analysis |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | baseline-selection | Select appropriate baselines for experimental comparison | | environment-specification | SOP: define complete experiment environment specification | | metric-specification | Define experiment metrics and significance standards | | sample-size-estimation | SOP: power analysis and required experiment count estimation | | seed-protocol-design | SOP: design random seed strategy for reproducibility |
<!-- END available-tables (generated) -->
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced…
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply.…
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's…
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the…
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to…
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall…