Skip to content
Automation
Skill

/assess-construct-validity

Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.

From plugin
de-anthropocentric-research-engine
499200 skills
Install
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill assess-construct-validity --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/assess-construct-validity

Context preview

The summary Claude sees to decide when to auto-load this skill.

Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.

SKILL.md

assess-construct-validity.SKILL.md
name: assess-construct-validity
description: "Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis."

assess-construct-validity

Purpose

Assess whether a benchmark or evaluation construct measures the claimed capability rather than content cues, confounds, or unrelated skill.

Input contract

required: [construct_claim, benchmark_specification, evaluation_records]
optional: [content_analysis, convergent_measures, discriminant_measures, confound_hypotheses]
constraints: [each validity judgment requires an observable indicator, comparison basis, and provenance]

Procedure

1. State the target construct and map benchmark tasks, labels, and metrics to its intended components. 2. Check content coverage and plausible construct-irrelevant cues against the benchmark specification. 3. Compare convergent and discriminant evidence where available, preserving missing comparisons. 4. Test confound hypotheses with controlled contrasts or artifact probes and record residual uncertainty.

If construct validity depends on whether the operationalization covers the intended domain, consider `map-coverage-space` as the next tactic.

Output contract

produces: [construct_map, content_validity_assessment, convergent_discriminant_evidence, confound_report, validity_judgment]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions]

Quality gates

  • Every validity claim cites a task, measure, contrast, or artifact observation.
  • Content coverage, convergence, discrimination, and confounds are reported separately.
  • A missing diagnostic is marked unresolved rather than treated as evidence of validity.

Failure and counterexamples

Do not infer construct validity from a high score, face validity, or agreement with another measure that shares the same confound.

Provenance map

  • `resolved: construct-validity-assessment`
Read more
Ships withde-anthropocentric-research-engine

The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.

Get the whole plugin
Stats
499
Stars
41
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
2h ago
Last commit
7mo ago
Created

Repo: yogsoth-ai/de-anthropocentric-research-engine