/benchmark-audit
Systematic quality assessment using BetterBench 46-criterion framework
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill benchmark-audit --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/benchmark-audit
Context preview
The summary Claude sees to decide when to auto-load this skill.
Systematic quality assessment using BetterBench 46-criterion framework
SKILL.md
benchmark-audit.SKILL.mdname: benchmark-audit
description: Systematic quality assessment using BetterBench 46-criterion framework
— 5 benchmarks, 30 papers, 40 web searches
dependencies:
tactics:
- artifact-detection
sops:
- benchmark-synthesis
- contamination-audit
- documentation-audit
- knowledge-acquisition-benchmark-inventory
- metric-decomposition
Benchmark Audit Strategy
Systematic quality assessment of AI/ML benchmarks using the BetterBench 46-criterion framework, Datasheets for Datasets standards, and established psychometric evaluation principles.
Purpose
Produce a structured quality report for each target benchmark covering: documentation completeness, construct validity indicators, statistical robustness, maintenance status, and known failure modes.
Budget
| Resource | Floor | Target | |----------|-------|--------| | Benchmarks audited | 3 | 5 | | Papers read | 20 | 30 | | Web searches | 25 | 40 |
State Ledger
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks audited | 0 | 5 | PENDING |
| Papers fetched | 0 | 30 | PENDING |
| Papers read | 0 | 20 | PENDING |
| Web searches | 0 | 40 | PENDING |
| Documentation audits complete | 0 | 5 | PENDING |
| Metric decompositions complete | 0 | 5 | PENDING |
| Contamination checks complete | 0 | 5 | PENDING |
| Synthesis reports produced | 0 | 5 | PENDING |
</HARD-GATE>
Cannot exit until 80% of all targets met.
Available Tactics
- **artifact-detection** — Probe for annotation artifacts and dataset shortcuts
Available SOPs
- **benchmark-inventory** — Identify target benchmarks in domain
- **metric-decomposition** — Decompose composite metrics into constituent signals
- **contamination-audit** — Detect train-test data leakage
- **documentation-audit** — Assess documentation completeness (BetterBench/Datasheets)
- **benchmark-synthesis** — Produce final structured audit report
Execution Guidance
1. **Inventory Phase**: Use benchmark-inventory to identify 5 benchmarks in target domain 2. **Per-Benchmark Loop** (repeat for each benchmark): a. Gather benchmark paper, documentation, leaderboard via web searches b. Run documentation-audit against BetterBench 46 criteria c. Run metric-decomposition on primary metric(s) d. Run contamination-audit checking known training corpora e. Run artifact-detection tactic if annotation-based benchmark f. Collect findings into per-benchmark report 3. **Synthesis Phase**: Run benchmark-synthesis to produce cross-benchmark comparison
Output Format
benchmark_audit:
benchmark_name: string
version: string
betterbench_score: float # 0-1, proportion of 46 criteria met
documentation_grade: A|B|C|D|F
metric_analysis:
primary_metric: string
ceiling_effects: boolean
polarity_issues: list
contamination_risk: low|medium|high|critical
artifact_risk: low|medium|high
maintenance_status: active|stale|abandoned
key_findings: list[string]
recommendations: list[string]<!-- BEGIN available-tables (generated) -->
Available Tactics
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | artifact-detection | Detect annotation artifacts and shortcuts in benchmarks |
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | benchmark-synthesis | Produce final structured audit report | | contamination-audit | Detect train-test data leakage and memorization artifacts | | documentation-audit | Assess documentation completeness against BetterBench/Datasheets standards | | knowledge-acquisition-benchmark-inventory | Identify and catalog all relevant benchmarks in target domain | | metric-decomposition | Decompose composite metrics into constituent signals, analyze polarity and ceiling effects |
<!-- END available-tables (generated) -->
Read more
name: benchmark-audit description: Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches dependencies: tactics: - artifact-detection sops: - benchmark-synthesis - contamination-audit - documentation-audit - knowledge-acquisition-benchmark-inventory - metric-decomposition
Benchmark Audit Strategy
Systematic quality assessment of AI/ML benchmarks using the BetterBench 46-criterion framework, Datasheets for Datasets standards, and established psychometric evaluation principles.
Purpose
Produce a structured quality report for each target benchmark covering: documentation completeness, construct validity indicators, statistical robustness, maintenance status, and known failure modes.
Budget
| Resource | Floor | Target | |----------|-------|--------| | Benchmarks audited | 3 | 5 | | Papers read | 20 | 30 | | Web searches | 25 | 40 |
State Ledger
<HARD-GATE> | Metric | Current | Target | Status | |--------|---------|--------|--------| | Benchmarks audited | 0 | 5 | PENDING | | Papers fetched | 0 | 30 | PENDING | | Papers read | 0 | 20 | PENDING | | Web searches | 0 | 40 | PENDING | | Documentation audits complete | 0 | 5 | PENDING | | Metric decompositions complete | 0 | 5 | PENDING | | Contamination checks complete | 0 | 5 | PENDING | | Synthesis reports produced | 0 | 5 | PENDING | </HARD-GATE>
Cannot exit until 80% of all targets met.
Available Tactics
- **artifact-detection** — Probe for annotation artifacts and dataset shortcuts
Available SOPs
- **benchmark-inventory** — Identify target benchmarks in domain
- **metric-decomposition** — Decompose composite metrics into constituent signals
- **contamination-audit** — Detect train-test data leakage
- **documentation-audit** — Assess documentation completeness (BetterBench/Datasheets)
- **benchmark-synthesis** — Produce final structured audit report
Execution Guidance
1. **Inventory Phase**: Use benchmark-inventory to identify 5 benchmarks in target domain 2. **Per-Benchmark Loop** (repeat for each benchmark): a. Gather benchmark paper, documentation, leaderboard via web searches b. Run documentation-audit against BetterBench 46 criteria c. Run metric-decomposition on primary metric(s) d. Run contamination-audit checking known training corpora e. Run artifact-detection tactic if annotation-based benchmark f. Collect findings into per-benchmark report 3. **Synthesis Phase**: Run benchmark-synthesis to produce cross-benchmark comparison
Output Format
benchmark_audit:
benchmark_name: string
version: string
betterbench_score: float # 0-1, proportion of 46 criteria met
documentation_grade: A|B|C|D|F
metric_analysis:
primary_metric: string
ceiling_effects: boolean
polarity_issues: list
contamination_risk: low|medium|high|critical
artifact_risk: low|medium|high
maintenance_status: active|stale|abandoned
key_findings: list[string]
recommendations: list[string]<!-- BEGIN available-tables (generated) -->
Available Tactics
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | artifact-detection | Detect annotation artifacts and shortcuts in benchmarks |
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | benchmark-synthesis | Produce final structured audit report | | contamination-audit | Detect train-test data leakage and memorization artifacts | | documentation-audit | Assess documentation completeness against BetterBench/Datasheets standards | | knowledge-acquisition-benchmark-inventory | Identify and catalog all relevant benchmarks in target domain | | metric-decomposition | Decompose composite metrics into constituent signals, analyze polarity and ceiling effects |
<!-- END available-tables (generated) -->
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Other skills on de-anthropocentric-research-engine.
- /formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced block in your reply. Do not execute the research.
Open skill - /formated-specs
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply. Replaces the generic spec-writing step.
Open skill - /injection-fidelity
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
Open skill - /ladder-quality-order
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
Open skill - /optimization-loop
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the
Open skill - /acu-nugget-recall
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.
Open skill

