formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced…
Aggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU, or V_strict/A_strict (+ run-level ranking, with an explicit per-topic-unreliability caveat) for Nugget. Use this as the final step of the atomic-unit chain,
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill atomic-unit-recall-aggregate --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/atomic-unit-recall-aggregateContext preview
The summary Claude sees to decide when to auto-load this skill.
Aggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU, or V_strict/A_strict (+ run-level ranking, with an explicit per-topic-unreliability caveat) for Nugget. Use this as the final step of the atomic-unit chain,
name: atomic-unit-recall-aggregate
description: Aggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU, or V_strict/A_strict (+ run-level ranking, with an explicit per-topic-unreliability caveat) for Nugget. Use this as the final step of the atomic-unit chain, after atomic-unit-matching; this SOP's existence closes a gap the original pipeline design was missing — without it, per-unit match judgments were never actually summed into the score the source methodologies report.
version: 1.0.0
category: paper-reading
type: sop
execution: subagent
prompt: ./prompt.md
input: 'match_results (list of {unit_text, judgment}), atomic_units (list of {text, importance})'
output: 'recall_score (float) OR {v_strict, a_strict, run_rank} depending on which method''s judgments were received'
dependencies:
sops:
- spawn-agentFinal aggregation step of the atomic-unit chain — normalized recall (ACU) or V_strict/A_strict + run-level ranking (Nugget). Added specifically to fix coverage-audit finding S5: the original graph's atomic-unit chain was a dead end at matching, with no node computing the actual reported score.
Subagent — spawned via spawn-agent skill.
Nugget-style per-topic scores are documented as unreliable (Kendall τ=0.297–0.539) — only run-level aggregation across multiple candidates is trustworthy (τ=0.887). This SOP must carry that caveat forward in its output whenever it runs Nugget-style aggregation on a single candidate; do not silence it for a cleaner-looking report.
<!-- BEGIN available-tables (generated) -->
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. |
<!-- END available-tables (generated) -->
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced…
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply.…
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's…
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the…
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to…
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall…