Skip to content
Content
Skill

/survey-methodology

Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative +

From plugin
luxas
68710 skills14 agents
Install
$ npx -y skills add Muuuun/luxas --skill survey-methodology --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/survey-methodology

Context preview

The summary Claude sees to decide when to auto-load this skill.

Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative +

SKILL.md

survey-methodology.SKILL.md
name: survey-methodology
description: Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative + Cochrane SRs, CS/ML/AI, math/Acta Numerica, earth/environment, economics/JEL, engineering/Annual Reviews, plus PRISMA/GRADE/Cochrane protocol literature. Empirical A-grade rate by domain ranges 7% (biology narrative) to 86% (math/Acta Numerica); the discriminator is structural, not stylistic. **Read this BEFORE writing notes/plan.md for any survey-style RESEARCH.md.**
compatibility: Pure prompt skill. References existing tools (spawn, escalate_authority_bound, experiment_reviewer, compile_latex).
allowed-tools: Read, Edit, Write, Glob, Grep

Survey Methodology Skill

The default failure mode of an autonomous-agent survey is **paper-trust**: read N papers, organize claims into a taxonomy, ship a prose digest. The output passes type-check (it looks like a survey) but fails verification (none of the cited numbers checked, contradictions not adjudicated, code not opened, negative space not bounded). This produces **B-grade** output.

Across ~240 reviews from 2024-2026, A-grade reviews share one structural discriminator:

> **Removing the new taxonomy from an A-grade survey leaves a contribution. > Removing it from a B-grade survey leaves nothing.**

Empirical A-rate by domain (with our wave-1 + wave-2 evidence base):

| Domain | A-rate | Modal A-pattern | |---|---|---| | Math (Acta Numerica / Bull AMS / SIAM Review / Probab Surv) | ~86% | Re-derivation in unified notation; new short proofs | | Economics (JEL / Annu Rev Econ / Handbook) | ~80% | Author re-estimation on harmonized data; "stylized-fact tables" | | Engineering (Annu Rev Control/BME, PECS, ARHT) | ~73% | Author re-simulation; harmonized device spec sheets | | Physics (RMP / Living Reviews / Annu Rev Cond Matt) | ~70% | Re-derivation + cross-paper number table; per-edition updates | | Chemistry/materials (Chem Rev / Chem Soc Rev / Annu Rev Phys Chem) | ~60% | Cross-paper benchmark table; Tutorial Review structured-closing | | Earth/environment (Rev Geophys / Annu Rev Earth Planet Sci / NRE&E) | ~40% | Narrative-with-embedded-re-analysis of observational data | | CS/ML/AI surveys (arXiv survey papers) | **~13%** | Bounded corpus + author benchmarks (BetterBench template) | | Biology narrative (Nature Reviews / Annu Rev Bio / Cell / Trends) | **~7%** | Almost never — venue norm is conceptual synthesis |

Cochrane / BMJ / Lancet SRs are 100% PRISMA-compliant by editorial policy but item-level adherence is asymmetric: ~75% of Cochrane abstracts use GRADE, but only ~7.5% of nominally compliant SRs across journals do full certainty + reporting-bias assessment. **The PRISMA label is not the substance** — verify item-by-item.

Two key empirical insights from the corpus:

1. **A-grade is topic-determined, not author-determined.** Surveys of *open artifacts* (open-source models, public conference proceedings, public datasets) admit A-grade execution. Surveys of *capabilities reported by closed systems* (RLHF/alignment, frontier-model agents, healthcare LLMs, industry-disclosed tools like Aletheia) are structurally trapped at B because the survey author cannot independently re-execute cited results.

2. **Disagreement-handling is a near-universal blind spot.** 0/31 CS surveys, ~12/30 biology reviews and ~9/30 physics reviews fence-sit on contradictions. Even A-grade work routinely fails this dimension. **It is the cleanest novelty axis the agent can exploit.**

When to use this skill

Trigger when RESEARCH.md uses: *survey, review, overview, landscape, state of the art, comparative analysis, taxonomy, benchmark of benchmarks, perspective.*

Skip for primary-research projects (single experiment + paper) — those use the standard experiment / experiment_reviewer pattern directly.

Step 1 — Pick the review type explicitly

Default-narrative is the modal mistake. An autonomous agent has no editorial-gatekeeping defense, so it inherits all narrative-review failure modes (cherry-picking, confirmation bias, irreproducibility) without the defenses. **Default to PRISMA-ScR-grade documentation at minimum.**

Choose one and commit it in `notes/scope.md` before any literature load:

| Type | When to choose | Required protocol | |---|---|---| | **Audit / benchmark survey** | Field has many primary systems with reported numbers; Q is "do the claims hold?" Most CS/ML/AI SOTA survey work falls here. | BetterBench-style: bounded N, criteria list, ≥2 raters, *count* don't gesture (Reuel/Balloccu template) | | **Scoping review** | Map breadth of a heterogeneous emerging field; decide whether full SR is warranted | PRISMA-ScR (Tricco 2018), 20 items, 5-stage Arksey-O'Malley. **No quality appraisal of included sources.** | | **Systematic review** | Bounded answerable question, evidence is appraisable | PRISMA 2020 (27 items) + RoB 2 / ROBINS-I + GRADE + PROSPERO registration | | **Umbrella review** | Synthesize multiple existing SRs on a related question | AMSTAR 2 for included reviews + handle SR overlap | | **Critical narrative review** | Domain conceptual synthesis where adjudication matters more than coverage (RMP-style theoretical recap; Annu Rev Phys Chem) | Greenhalgh: explicit interpreter positioning + explicit selection logic + explicit acknowledgement of evidence not selected. **No paper-trust.** | | **Narrative-with-embedded-re-analysis** | Earth/environment / climate where review value-add is reprocessing observational datasets | `notes/datasets.md` provenance + reproducible reprocessing pipeline | | **Theoretical-unification survey** | Math/theoretical review where unifying object is the contribution (Acta Numerica template) | Re-derivation in unified notation; new short proofs of known results; compet

Read more
Ships withluxas

An autonomous research colleague — from a question to a compiled manuscript, while you sleep.

Get the whole plugin

Other skills on luxas.