/artifact-detection
Detect annotation artifacts and shortcuts in benchmarks
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill artifact-detection --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/artifact-detection
Context preview
The summary Claude sees to decide when to auto-load this skill.
Detect annotation artifacts and shortcuts in benchmarks
SKILL.md
artifact-detection.SKILL.mdname: artifact-detection
description: Detect annotation artifacts and shortcuts in benchmarks
execution: tactic
Artifact Detection Tactic
Systematically probe benchmarks for annotation artifacts, dataset shortcuts, and spurious correlations that allow models to achieve high scores without the intended capability.
Stages
Stage 1: Hypothesis-Only Baseline Test
Search literature for evidence that partial-input baselines achieve unexpectedly high performance:
- Hypothesis-only baselines (NLI without premise)
- Question-only baselines (QA without context)
- Label-word frequency baselines
- Majority-class and surface-pattern baselines
**Search queries**: "[benchmark] annotation artifacts", "[benchmark] hypothesis only", "[benchmark] spurious correlations", "[benchmark] dataset bias"
If published partial-input results exist, record performance gap between partial and full input. Gap < 10 points above random indicates severe artifacts.
Stage 2: Contrast Set Construction
Identify whether contrast sets or adversarial evaluations exist:
- Search for "[benchmark] contrast sets", "[benchmark] adversarial examples"
- Check if CheckList-style behavioral tests have been applied
- Look for counterfactual data augmentation studies
Record performance drops on contrast sets. Drops > 20 points indicate reliance on surface patterns.
Stage 3: Format Manipulation Probes
Search for evidence of format sensitivity:
- Prompt template sensitivity studies
- Label name/ordering effects
- Verbalization effects in classification
- Input length correlations with labels
Record whether minor format changes cause disproportionate score changes.
Stage 4: Conclusion Synthesis
Aggregate evidence into artifact severity assessment:
| Severity | Criteria | |----------|----------| | Critical | Partial-input baseline within 5 points of full model | | High | Contrast set drop >20 points OR format sensitivity >10 points | | Medium | Known artifacts documented but partial mitigations exist | | Low | Minor artifacts, full-input still required for high performance | | None | No evidence of artifacts (may indicate insufficient probing) |
Output
artifact_report:
benchmark: string
overall_severity: critical|high|medium|low|none
partial_input_baselines:
- input_type: string # e.g., "hypothesis only"
performance: float
full_model_performance: float
gap: float
source: string
contrast_set_results:
- contrast_set: string
original_performance: float
contrast_performance: float
drop: float
source: string
format_sensitivity:
- manipulation: string
score_range: string
source: string
shortcuts_identified:
- shortcut: string
mechanism: string
exploitability: high|medium|low
evidence_completeness: thorough|partial|minimalYield Report
| Metric | Minimum | |--------|---------| | Literature sources checked | 5 | | Artifact categories probed | 3 | | Evidence items collected | 4 | | Severity classification produced | 1 |
Read more
name: artifact-detection description: Detect annotation artifacts and shortcuts in benchmarks execution: tactic
Artifact Detection Tactic
Systematically probe benchmarks for annotation artifacts, dataset shortcuts, and spurious correlations that allow models to achieve high scores without the intended capability.
Stages
Stage 1: Hypothesis-Only Baseline Test
Search literature for evidence that partial-input baselines achieve unexpectedly high performance:
- Hypothesis-only baselines (NLI without premise)
- Question-only baselines (QA without context)
- Label-word frequency baselines
- Majority-class and surface-pattern baselines
**Search queries**: "[benchmark] annotation artifacts", "[benchmark] hypothesis only", "[benchmark] spurious correlations", "[benchmark] dataset bias"
If published partial-input results exist, record performance gap between partial and full input. Gap < 10 points above random indicates severe artifacts.
Stage 2: Contrast Set Construction
Identify whether contrast sets or adversarial evaluations exist:
- Search for "[benchmark] contrast sets", "[benchmark] adversarial examples"
- Check if CheckList-style behavioral tests have been applied
- Look for counterfactual data augmentation studies
Record performance drops on contrast sets. Drops > 20 points indicate reliance on surface patterns.
Stage 3: Format Manipulation Probes
Search for evidence of format sensitivity:
- Prompt template sensitivity studies
- Label name/ordering effects
- Verbalization effects in classification
- Input length correlations with labels
Record whether minor format changes cause disproportionate score changes.
Stage 4: Conclusion Synthesis
Aggregate evidence into artifact severity assessment:
| Severity | Criteria | |----------|----------| | Critical | Partial-input baseline within 5 points of full model | | High | Contrast set drop >20 points OR format sensitivity >10 points | | Medium | Known artifacts documented but partial mitigations exist | | Low | Minor artifacts, full-input still required for high performance | | None | No evidence of artifacts (may indicate insufficient probing) |
Output
artifact_report:
benchmark: string
overall_severity: critical|high|medium|low|none
partial_input_baselines:
- input_type: string # e.g., "hypothesis only"
performance: float
full_model_performance: float
gap: float
source: string
contrast_set_results:
- contrast_set: string
original_performance: float
contrast_performance: float
drop: float
source: string
format_sensitivity:
- manipulation: string
score_range: string
source: string
shortcuts_identified:
- shortcut: string
mechanism: string
exploitability: high|medium|low
evidence_completeness: thorough|partial|minimalYield Report
| Metric | Minimum | |--------|---------| | Literature sources checked | 5 | | Artifact categories probed | 3 | | Evidence items collected | 4 | | Severity classification produced | 1 |
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Other skills on de-anthropocentric-research-engine.
- /formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced block in your reply. Do not execute the research.
Open skill - /formated-specs
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply. Replaces the generic spec-writing step.
Open skill - /injection-fidelity
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
Open skill - /ladder-quality-order
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
Open skill - /optimization-loop
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the
Open skill - /acu-nugget-recall
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.
Open skill

