abstract-structure
Remove domain surface details to expose transferable relational/mechanistic structure at a chosen abstraction level.
Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill audit-benchmark-validity --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/audit-benchmark-validityContext preview
The summary Claude sees to decide when to auto-load this skill.
Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.
name: audit-benchmark-validity description: "Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift."
Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.
required: [benchmark_spec, task_definition, metric_definition, evaluation_records] optional: [leaderboard_history, protocol_versions, coverage_target] constraints: [claims must be linked to benchmark evidence]
Do not perform called SOP operations inline; each loaded SOP owns its contract and thresholds.
1. You MUST load skill `inventory-reference-items` to inventory benchmark components and versions. You MUST load skill `decompose-evaluation-metric` to decompose the evaluation metric. You MUST load skill `assess-construct-validity` to assess the benchmark construct. You MUST load skill `extract-evaluation-protocol` to reconstruct protocol versions. 2. Select validity, contamination, saturation, coverage, protocol-forensics, or evaluation-comparison mode. 3. You MUST load skill `audit-data-contamination` to audit contamination. You MUST load skill `map-coverage-space` to map data and task coverage. You MUST load skill `analyze-leaderboard-dynamics` to analyze saturation and temporal dynamics. You MUST load skill `compare-evaluation-protocols` to compare protocol and metric variants. You MUST load skill `probe-benchmark-artifact` to probe artifacts and shortcut paths. 4. You MUST load skill `audit-reporting-quality` to audit reporting quality and return the validity verdict with threats, evidence, and required repairs. If the audit exposes systematic coverage gaps, consider `coverage-white-space-search`. If the benchmark or metric may share assumptions with the tested system, consider `audit-validator-independence`.
produces: [validity_verdict, threat_register, contamination_findings, coverage_map, protocol_drift_report, repair_actions] delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions, recommended_jumps]
Reject "valid" when benchmark artifact probes are absent, contamination is unknown but ignored, or leaderboard gains cannot be separated from protocol drift. A high score is not evidence of construct validity by itself.
7 architecture `old` entries: archaeology, audit, saturation, validity probing, coverage mapping, protocol forensics, evaluation comparison. Provider-specific retrieval compressed.
Append benchmark version, construct claims, probes, contamination evidence, coverage gaps, protocol diffs, verdict, and repair decisions.
| source | source line | kind | source criterion | |---|---:|---|---| | benchmark-archaeology | 39 | numeric-table | \\| benchmark-audit \\| Systematic quality assessment using BetterBench 46-criterion framework \\| | | benchmark-archaeology | 72 | numeric-table | \\| benchmark-audit \\| 5 \\| 30 \\| 40 \\| | | benchmark-archaeology | 73 | numeric-table | \\| saturation-analysis \\| 15 \\| 50 \\| 60 \\| | | benchmark-archaeology | 74 | numeric-table | \\| validity-probing \\| 3 \\| 40 \\| 30 \\| | | benchmark-archaeology | 75 | numeric-table | \\| coverage-mapping \\| 20 \\| 30 \\| 50 \\| | | benchmark-archaeology | 76 | numeric-table | \\| protocol-forensics \\| 5 \\| 60 \\| 30 \\| | | benchmark-archaeology | 77 | numeric-table | \\| **Total** \\| **48** \\| **210** \\| **210** \\| | | benchmark-audit | 28 | numeric-table | \\| Benchmarks audited \\| 3 \\| 5 \\| | | benchmark-audit | 29 | numeric-table | \\| Papers read \\| 20 \\| 30 \\| | | benchmark-audit | 30 | numeric-table | \\| Web searches \\| 25 \\| 40 \\| | | benchmark-audit | 35 | textual | <HARD-GATE> | | benchmark-audit | 38 | numeric-table | \\| Benchmarks audited \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 39 | numeric-table | \\| Papers fetched \\| 0 \\| 30 \\| PENDING \\| | | benchmark-audit | 40 | numeric-table | \\| Papers read \\| 0 \\| 20 \\| PENDING \\| | | benchmark-audit | 41 | numeric-table | \\| Web searches \\| 0 \\| 40 \\| PENDING \\| | | benchmark-audit | 42 | numeric-table | \\| Documentation audits complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 43 | numeric-table | \\| Metric decompositions complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 44 | numeric-table | \\| Contamination checks complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 45 | numeric-table | \\| Synthesis reports produced \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 46 | textual | </HARD-GATE> | | benchmark-audit | 49 | numeric | Cannot exit until 80% of all targets met. | | benchmark-audit | 81 | numeric | betterbench_score: float # 0-1, proportion of 46 criteria me
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Remove domain surface details to expose transferable relational/mechanistic structure at a chosen abstraction level.
Evaluate competing arguments against stated criteria and produce a reasoned verdict with uncertainty.
Move a scientific object up/down in abstraction or narrow/broaden selected scope dimensions (population, mechanism, context, outcome, timeframe, system…
Run structured attack/defense/adjudication over a claim, candidate, criterion set, or current winner. Perspective, target, escalation depth,…
Aggregate criterion or comparison results into an ordered recommendation under an explicit rule.
Abstract relational structure from source domains, map it to the target, validate depth, and instantiate transferable mechanisms.