/ladder-quality-order
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill ladder-quality-order --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/ladder-quality-order
Context preview
The summary Claude sees to decide when to auto-load this skill.
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
SKILL.md
ladder-quality-order.SKILL.mdname: ladder-quality-order
description: Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
ladder-quality-order (loss-2)
You rank ONE topic's 6 research-design samples (each a research_graph + research_result pair) by quality. The samples arrive SHUFFLED and anonymous — you see 6 positions (0–5), never their true rung id or config. You judge only on the D1–D5 standard:
- D1 meaningfulness — is the research question real and worth asking?
- D2 skill-research value — does the design advance skill/methodology research?
- D3 use-to-DARE — is it usable by the DARE engine?
- D4 respects the 4-layer architecture (campaign → strategy → tactic → sop)?
- D5 prerequisites — are the stated prerequisites sound and met?
Judge only on the D1–D5 standard above; never on academic-publication criteria of any kind. You never see any quality-check list.
Pairwise mechanism
You will be asked to compare two positions at a time. For each pair `(i, j)` decide the `winner` (the higher-quality position) and give a one-line `reason` grounded in D1–D5. Do not assign absolute scores — only pick a winner per pair. The graph is structure-aware context; read it holistically, do not run any checklist over it.
The harness enumerates all 15 pairs (i<j over 6 positions), Copeland-aggregates your winners into an induced order, un-shuffles to true ids, and computes Kendall τ against the intended order id0 > id1 > … > id5 (id0 = highest quality). You only emit `{winner, reason}` per pair.
Endpoint separation
You will also be asked, K independent times, to compare the two extreme samples (the harness picks them and presents them as just two options, **A** and **B**). Return `{"winner": "A" | "B"}` — exactly the label of the higher-quality one. Judge each call independently and honestly; do not try to be consistent with a previous call you don't remember. (This is a two-way A/B label, distinct from the position integers used in the pairwise rank above.)
Confound flat-check (when present)
If the topic carries a same-substance / different-framing triplet, rank it first. The order must NOT change with framing alone (buzzword vs neutral wording is not a quality difference under D1–D5). If your order tracks framing, say so in the reason — the harness will treat this topic's ladder as untrustworthy.
What the harness writes (you do not compute these)
The harness assembles `loss2.json`: `tau`, `monotonicity_pass` (τ ≥ the τ line AND no adjacent endpoint inversion), `endpoint_separation_pass` (endpoint majority), `rigor_floor_flag` (endpoints a near-tie — a possible genuine quality floor, NOT a tuning bug), and `pairwise_log` (your winners + reasons, un-shuffled to true ids). Your only job: honest per-pair winners and D1–D5 reasons.
Read more
name: ladder-quality-order description: Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
ladder-quality-order (loss-2)
You rank ONE topic's 6 research-design samples (each a research_graph + research_result pair) by quality. The samples arrive SHUFFLED and anonymous — you see 6 positions (0–5), never their true rung id or config. You judge only on the D1–D5 standard:
- D1 meaningfulness — is the research question real and worth asking?
- D2 skill-research value — does the design advance skill/methodology research?
- D3 use-to-DARE — is it usable by the DARE engine?
- D4 respects the 4-layer architecture (campaign → strategy → tactic → sop)?
- D5 prerequisites — are the stated prerequisites sound and met?
Judge only on the D1–D5 standard above; never on academic-publication criteria of any kind. You never see any quality-check list.
Pairwise mechanism
You will be asked to compare two positions at a time. For each pair `(i, j)` decide the `winner` (the higher-quality position) and give a one-line `reason` grounded in D1–D5. Do not assign absolute scores — only pick a winner per pair. The graph is structure-aware context; read it holistically, do not run any checklist over it.
The harness enumerates all 15 pairs (i<j over 6 positions), Copeland-aggregates your winners into an induced order, un-shuffles to true ids, and computes Kendall τ against the intended order id0 > id1 > … > id5 (id0 = highest quality). You only emit `{winner, reason}` per pair.
Endpoint separation
You will also be asked, K independent times, to compare the two extreme samples (the harness picks them and presents them as just two options, **A** and **B**). Return `{"winner": "A" | "B"}` — exactly the label of the higher-quality one. Judge each call independently and honestly; do not try to be consistent with a previous call you don't remember. (This is a two-way A/B label, distinct from the position integers used in the pairwise rank above.)
Confound flat-check (when present)
If the topic carries a same-substance / different-framing triplet, rank it first. The order must NOT change with framing alone (buzzword vs neutral wording is not a quality difference under D1–D5). If your order tracks framing, say so in the reason — the harness will treat this topic's ladder as untrustworthy.
What the harness writes (you do not compute these)
The harness assembles `loss2.json`: `tau`, `monotonicity_pass` (τ ≥ the τ line AND no adjacent endpoint inversion), `endpoint_separation_pass` (endpoint majority), `rigor_floor_flag` (endpoints a near-tie — a possible genuine quality floor, NOT a tuning bug), and `pairwise_log` (your winners + reasons, un-shuffled to true ids). Your only job: honest per-pair winners and D1–D5 reasons.
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Other skills on de-anthropocentric-research-engine.
- /formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced block in your reply. Do not execute the research.
Open skill - /formated-specs
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply. Replaces the generic spec-writing step.
Open skill - /injection-fidelity
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
Open skill - /optimization-loop
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the
Open skill - /acu-nugget-recall
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.
Open skill - /argumentative-zoning
Tactic: Label every sentence of one paper with its rhetorical role using Argumentative Zoning. Use when fixed rhetorical labels and cross-paper alignment matter.
Open skill

