abstract-structure
Remove domain surface details to expose transferable relational/mechanistic structure at a chosen abstraction level.
Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill assess-construct-validity --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/assess-construct-validityContext preview
The summary Claude sees to decide when to auto-load this skill.
Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.
name: assess-construct-validity description: "Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis."
Assess whether a benchmark or evaluation construct measures the claimed capability rather than content cues, confounds, or unrelated skill.
required: [construct_claim, benchmark_specification, evaluation_records] optional: [content_analysis, convergent_measures, discriminant_measures, confound_hypotheses] constraints: [each validity judgment requires an observable indicator, comparison basis, and provenance]
1. State the target construct and map benchmark tasks, labels, and metrics to its intended components. 2. Check content coverage and plausible construct-irrelevant cues against the benchmark specification. 3. Compare convergent and discriminant evidence where available, preserving missing comparisons. 4. Test confound hypotheses with controlled contrasts or artifact probes and record residual uncertainty.
If construct validity depends on whether the operationalization covers the intended domain, consider `map-coverage-space` as the next tactic.
produces: [construct_map, content_validity_assessment, convergent_discriminant_evidence, confound_report, validity_judgment] delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions]
Do not infer construct validity from a high score, face validity, or agreement with another measure that shares the same confound.
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Remove domain surface details to expose transferable relational/mechanistic structure at a chosen abstraction level.
Evaluate competing arguments against stated criteria and produce a reasoned verdict with uncertainty.
Move a scientific object up/down in abstraction or narrow/broaden selected scope dimensions (population, mechanism, context, outcome, timeframe, system…
Run structured attack/defense/adjudication over a claim, candidate, criterion set, or current winner. Perspective, target, escalation depth,…
Aggregate criterion or comparison results into an ordered recommendation under an explicit rule.
Abstract relational structure from source domains, map it to the target, validate depth, and instantiate transferable mechanisms.