/collective-adjudication
Strategy for multi-judge ranking aggregation using Condorcet, Schulze,
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill collective-adjudication --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/collective-adjudication
Context preview
The summary Claude sees to decide when to auto-load this skill.
Strategy for multi-judge ranking aggregation using Condorcet, Schulze,
SKILL.md
collective-adjudication.SKILL.mdname: collective-adjudication
description: Strategy for multi-judge ranking aggregation using Condorcet, Schulze,
Borda, Kemeny-Young, and Copeland methods to produce consensus rankings from diverse
perspectives.
dependencies:
tactics:
- consistency-audit-loop
- multi-judge-aggregation
sops:
- ranking-synthesis
Collective Adjudication
Purpose
Aggregate rankings from multiple independent judges into a single consensus ranking. Handles disagreement detection, voting paradoxes, and produces transparent aggregation with disagreement maps.
When to use
- Multiple judges/evaluators available (≥3)
- LLM-as-judge with multiple prompting perspectives
- Committee decision-making requiring formal aggregation
- Need to identify and characterize disagreement patterns
Budget
| Resource | Allocation | |----------|-----------| | Judges/Perspectives | ≥3 independent evaluators | | Comparisons per judge | Complete or near-complete per judge | | Aggregation methods | ≥2 methods for robustness check | | Disagreement threshold | Flag pairs where judges disagree >40% |
State Ledger
candidates: []
perspectives: [] # judge identities/prompts
ballots: [] # [{judge, ranking: [...]}]
aggregation_results: {} # method → consensus_ranking
disagreement_map: {} # pair → {agreement_rate, split}
cycles: [] # Condorcet cycles if any
method: "" # schulze | borda | kemeny-young | copelandAvailable Tactics
- **multi-judge-aggregation** — collect ballots, aggregate, identify disagreement
- **consistency-audit-loop** — detect cycles in aggregated preferences
Available SOPs
- ballot-collection
- aggregation-method
- cycle-detection
- inconsistency-localization
- ranking-synthesis
Execution Guidance
1. Define perspectives (judge roles, prompting strategies) 2. Run ballot-collection to gather independent rankings 3. Run aggregation-method with primary method (Schulze recommended) 4. Run cycle-detection on aggregated pairwise matrix 5. If cycles exist, run inconsistency-localization 6. Cross-validate with secondary method (Borda or Copeland) 7. Produce final ranking with disagreement heatmap
Output Format
consensus_ranking:
- {rank: 1, candidate: "...", wins: 8, copeland_score: 0.95}
- {rank: 2, candidate: "...", wins: 7, copeland_score: 0.88}
method: schulze
judges: 5
condorcet_winner: "candidate_a" # or null if cycle
disagreement_hotspots:
- {pair: ["c", "d"], agreement: 0.4, split: "3:2"}
cross_validation: {borda_agreement: 0.92, copeland_agreement: 0.96}<!-- BEGIN available-tables (generated) -->
Available Tactics
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | consistency-audit-loop | Detect preference cycles, localize inconsistent judgments, request corrections, and recompute ratings until consistency threshold is met. | | multi-judge-aggregation | Collect independent rankings from multiple judges, aggregate using social choice methods, and identify disagreement hotspots. |
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | ranking-synthesis | Produce the final ranking artifact from converged ratings and consistency report. |
<!-- END available-tables (generated) -->
Read more
name: collective-adjudication description: Strategy for multi-judge ranking aggregation using Condorcet, Schulze, Borda, Kemeny-Young, and Copeland methods to produce consensus rankings from diverse perspectives. dependencies: tactics: - consistency-audit-loop - multi-judge-aggregation sops: - ranking-synthesis
Collective Adjudication
Purpose
Aggregate rankings from multiple independent judges into a single consensus ranking. Handles disagreement detection, voting paradoxes, and produces transparent aggregation with disagreement maps.
When to use
- Multiple judges/evaluators available (≥3)
- LLM-as-judge with multiple prompting perspectives
- Committee decision-making requiring formal aggregation
- Need to identify and characterize disagreement patterns
Budget
| Resource | Allocation | |----------|-----------| | Judges/Perspectives | ≥3 independent evaluators | | Comparisons per judge | Complete or near-complete per judge | | Aggregation methods | ≥2 methods for robustness check | | Disagreement threshold | Flag pairs where judges disagree >40% |
State Ledger
candidates: []
perspectives: [] # judge identities/prompts
ballots: [] # [{judge, ranking: [...]}]
aggregation_results: {} # method → consensus_ranking
disagreement_map: {} # pair → {agreement_rate, split}
cycles: [] # Condorcet cycles if any
method: "" # schulze | borda | kemeny-young | copelandAvailable Tactics
- **multi-judge-aggregation** — collect ballots, aggregate, identify disagreement
- **consistency-audit-loop** — detect cycles in aggregated preferences
Available SOPs
- ballot-collection
- aggregation-method
- cycle-detection
- inconsistency-localization
- ranking-synthesis
Execution Guidance
1. Define perspectives (judge roles, prompting strategies) 2. Run ballot-collection to gather independent rankings 3. Run aggregation-method with primary method (Schulze recommended) 4. Run cycle-detection on aggregated pairwise matrix 5. If cycles exist, run inconsistency-localization 6. Cross-validate with secondary method (Borda or Copeland) 7. Produce final ranking with disagreement heatmap
Output Format
consensus_ranking:
- {rank: 1, candidate: "...", wins: 8, copeland_score: 0.95}
- {rank: 2, candidate: "...", wins: 7, copeland_score: 0.88}
method: schulze
judges: 5
condorcet_winner: "candidate_a" # or null if cycle
disagreement_hotspots:
- {pair: ["c", "d"], agreement: 0.4, split: "3:2"}
cross_validation: {borda_agreement: 0.92, copeland_agreement: 0.96}<!-- BEGIN available-tables (generated) -->
Available Tactics
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | consistency-audit-loop | Detect preference cycles, localize inconsistent judgments, request corrections, and recompute ratings until consistency threshold is met. | | multi-judge-aggregation | Collect independent rankings from multiple judges, aggregate using social choice methods, and identify disagreement hotspots. |
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | ranking-synthesis | Produce the final ranking artifact from converged ratings and consistency report. |
<!-- END available-tables (generated) -->
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Other skills on de-anthropocentric-research-engine.
- /formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced block in your reply. Do not execute the research.
Open skill - /formated-specs
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply. Replaces the generic spec-writing step.
Open skill - /injection-fidelity
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
Open skill - /ladder-quality-order
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
Open skill - /optimization-loop
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the
Open skill - /acu-nugget-recall
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.
Open skill

