Skip to content
Automation
Skill

/audit-benchmark-validity

Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.

From plugin
de-anthropocentric-research-engine
499200 skills
Install
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill audit-benchmark-validity --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/audit-benchmark-validity

Context preview

The summary Claude sees to decide when to auto-load this skill.

Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.

SKILL.md

audit-benchmark-validity.SKILL.md
name: audit-benchmark-validity
description: "Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift."

audit-benchmark-validity

Purpose

Treat benchmarks as scientific measurement instruments: audit construct validity, contamination, metric pathology, coverage, leaderboard dynamics, and protocol drift.

Input contract

required: [benchmark_spec, task_definition, metric_definition, evaluation_records]
optional: [leaderboard_history, protocol_versions, coverage_target]
constraints: [claims must be linked to benchmark evidence]

Execution protocol

Do not perform called SOP operations inline; each loaded SOP owns its contract and thresholds.

1. You MUST load skill `inventory-reference-items` to inventory benchmark components and versions. You MUST load skill `decompose-evaluation-metric` to decompose the evaluation metric. You MUST load skill `assess-construct-validity` to assess the benchmark construct. You MUST load skill `extract-evaluation-protocol` to reconstruct protocol versions. 2. Select validity, contamination, saturation, coverage, protocol-forensics, or evaluation-comparison mode. 3. You MUST load skill `audit-data-contamination` to audit contamination. You MUST load skill `map-coverage-space` to map data and task coverage. You MUST load skill `analyze-leaderboard-dynamics` to analyze saturation and temporal dynamics. You MUST load skill `compare-evaluation-protocols` to compare protocol and metric variants. You MUST load skill `probe-benchmark-artifact` to probe artifacts and shortcut paths. 4. You MUST load skill `audit-reporting-quality` to audit reporting quality and return the validity verdict with threats, evidence, and required repairs. If the audit exposes systematic coverage gaps, consider `coverage-white-space-search`. If the benchmark or metric may share assumptions with the tested system, consider `audit-validator-independence`.

Output contract

produces: [validity_verdict, threat_register, contamination_findings, coverage_map, protocol_drift_report, repair_actions]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions, recommended_jumps]

Thresholds and quality gates

  • Every source HARD-GATE remains mandatory; no exit with an untested construct, contamination path, metric pathology, coverage claim, or protocol change.
  • Resource and sampling gates use relative benchmark/target/evidence coverage rather than fixed benchmark, paper, or web counts. Declare the eligible universe, record numerator, denominator, batch increment, stopping reason, and source references.
  • Saturation claims require an explicit stopping criterion and evidence that additional search/testing no longer changes the conclusion; report marginal information gain and saturation state.
  • BetterBench-style criterion lists remain content checklists. Their item count is not converted into a percentage; the audit reports criterion coverage ratio over the declared applicable set and an independent-source ratio.
  • Evaluation comparisons must state the controlled protocol difference and its expected impact.

Failure and counterexamples

Reject "valid" when benchmark artifact probes are absent, contamination is unknown but ignored, or leaderboard gains cannot be separated from protocol drift. A high score is not evidence of construct validity by itself.

Provenance map

7 architecture `old` entries: archaeology, audit, saturation, validity probing, coverage mapping, protocol forensics, evaluation comparison. Provider-specific retrieval compressed.

Legacy context checkpoint / Delta notes

Append benchmark version, construct claims, probes, contamination evidence, coverage gaps, protocol diffs, verdict, and repair decisions.

Preserved source criteria ledger

| source | source line | kind | source criterion | |---|---:|---|---| | benchmark-archaeology | 39 | numeric-table | \\| benchmark-audit \\| Systematic quality assessment using BetterBench 46-criterion framework \\| | | benchmark-archaeology | 72 | numeric-table | \\| benchmark-audit \\| 5 \\| 30 \\| 40 \\| | | benchmark-archaeology | 73 | numeric-table | \\| saturation-analysis \\| 15 \\| 50 \\| 60 \\| | | benchmark-archaeology | 74 | numeric-table | \\| validity-probing \\| 3 \\| 40 \\| 30 \\| | | benchmark-archaeology | 75 | numeric-table | \\| coverage-mapping \\| 20 \\| 30 \\| 50 \\| | | benchmark-archaeology | 76 | numeric-table | \\| protocol-forensics \\| 5 \\| 60 \\| 30 \\| | | benchmark-archaeology | 77 | numeric-table | \\| **Total** \\| **48** \\| **210** \\| **210** \\| | | benchmark-audit | 28 | numeric-table | \\| Benchmarks audited \\| 3 \\| 5 \\| | | benchmark-audit | 29 | numeric-table | \\| Papers read \\| 20 \\| 30 \\| | | benchmark-audit | 30 | numeric-table | \\| Web searches \\| 25 \\| 40 \\| | | benchmark-audit | 35 | textual | <HARD-GATE> | | benchmark-audit | 38 | numeric-table | \\| Benchmarks audited \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 39 | numeric-table | \\| Papers fetched \\| 0 \\| 30 \\| PENDING \\| | | benchmark-audit | 40 | numeric-table | \\| Papers read \\| 0 \\| 20 \\| PENDING \\| | | benchmark-audit | 41 | numeric-table | \\| Web searches \\| 0 \\| 40 \\| PENDING \\| | | benchmark-audit | 42 | numeric-table | \\| Documentation audits complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 43 | numeric-table | \\| Metric decompositions complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 44 | numeric-table | \\| Contamination checks complete \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 45 | numeric-table | \\| Synthesis reports produced \\| 0 \\| 5 \\| PENDING \\| | | benchmark-audit | 46 | textual | </HARD-GATE> | | benchmark-audit | 49 | numeric | Cannot exit until 80% of all targets met. | | benchmark-audit | 81 | numeric | betterbench_score: float # 0-1, proportion of 46 criteria me

Read more
Ships withde-anthropocentric-research-engine

The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.

Get the whole plugin
Stats
499
Stars
41
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
2h ago
Last commit
7mo ago
Created

Repo: yogsoth-ai/de-anthropocentric-research-engine