Skip to content
Automation
Skill

/comparison-design

Design fair comparison experiments against baselines and competing methods

From plugin
de-anthropocentric-research-engine
393200 skills
Install
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill comparison-design --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/comparison-design

Context preview

The summary Claude sees to decide when to auto-load this skill.

Design fair comparison experiments against baselines and competing methods

SKILL.md

comparison-design.SKILL.md
name: comparison-design
description: Design fair comparison experiments against baselines and competing methods
version: 1.0.0
category: experiment-execution
type: strategy
sops:
- baseline-selection
- metric-specification
- sample-size-estimation
- seed-protocol-design
- environment-specification
tactics:
- statistical-method-selection
- reproducibility-protocol
dependencies:
  sops:
  - baseline-selection
  - environment-specification
  - metric-specification
  - sample-size-estimation
  - seed-protocol-design
  tactics:
  - reproducibility-protocol
  - statistical-method-selection

Strategy: Comparison Design

**Question**: How much better is our method than the baseline?

Methodology

  • **Fair Comparison Protocol** (Bouthillier 2021): Control all confounds, same compute budget, same tuning effort.
  • **Multi-Baseline Comparison**: Compare against multiple baselines (SOTA, simple, ablated).
  • **Multi-Dataset Evaluation**: Test across diverse datasets to avoid dataset-specific overfitting.
  • **Bayesian Comparison** (Benavoli 2017): Posterior probability of superiority, not just p-values.
  • **Bootstrap/Permutation Tests**: Non-parametric significance without distributional assumptions.

Execution Flow

1. **baseline-selection** → Select appropriate baselines (SOTA, simple, oracle) 2. **metric-specification** → Define primary metric and secondary metrics 3. **sample-size-estimation** → Power analysis for detecting meaningful differences 4. **seed-protocol-design** → Ensure fair random initialization across methods 5. **environment-specification** → Lock environment to prevent confounds 6. **reproducibility-protocol** (tactic) → Ensure all results are reproducible 7. **statistical-method-selection** (tactic) → Choose Bayesian or frequentist comparison

Budget Gate

| Comparison Scope | Baselines | Datasets | Seeds | Min Runs | |-----------------|-----------|----------|-------|----------| | Minimal | 1 SOTA + 1 simple | 1 | 3 | 6 | | Standard | 2-3 baselines | 2-3 | 5 | 30-45 | | Comprehensive | 4+ baselines | 3-5 | 5-10 | 100+ | | Publication-ready | All relevant | 5+ | 10+ | 200+ |

<!-- BEGIN available-tables (generated) -->

Available Tactics

Optional, no fixed order; the final leaf is always a sop.

| Tactic | When to use | | --- | --- | | reproducibility-protocol | Ensure experiment reproducibility through systematic environment and seed control | | statistical-method-selection | Select appropriate statistical methods for experiment analysis |

Available SOPs

Optional, no fixed order; the final leaf is always a sop.

| SOP | When to use | | --- | --- | | baseline-selection | Select appropriate baselines for experimental comparison | | environment-specification | SOP: define complete experiment environment specification | | metric-specification | Define experiment metrics and significance standards | | sample-size-estimation | SOP: power analysis and required experiment count estimation | | seed-protocol-design | SOP: design random seed strategy for reproducibility |

<!-- END available-tables (generated) -->

Read more
Ships withde-anthropocentric-research-engine

The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.

Get the whole plugin
Stats
393
Stars
34
Forks
Active
Maintenance
HTML
Language
Apache-2.0
License
19h ago
Last commit
6mo ago
Created

Repo: yogsoth-ai/de-anthropocentric-research-engine