/survey-methodology
Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative +
$ npx -y skills add Muuuun/luxas --skill survey-methodology --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/survey-methodology
Context preview
The summary Claude sees to decide when to auto-load this skill.
Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative +
SKILL.md
survey-methodology.SKILL.mdname: survey-methodology
description: Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative + Cochrane SRs, CS/ML/AI, math/Acta Numerica, earth/environment, economics/JEL, engineering/Annual Reviews, plus PRISMA/GRADE/Cochrane protocol literature. Empirical A-grade rate by domain ranges 7% (biology narrative) to 86% (math/Acta Numerica); the discriminator is structural, not stylistic. **Read this BEFORE writing notes/plan.md for any survey-style RESEARCH.md.**
compatibility: Pure prompt skill. References existing tools (spawn, escalate_authority_bound, experiment_reviewer, compile_latex).
allowed-tools: Read, Edit, Write, Glob, Grep
Survey Methodology Skill
The default failure mode of an autonomous-agent survey is **paper-trust**: read N papers, organize claims into a taxonomy, ship a prose digest. The output passes type-check (it looks like a survey) but fails verification (none of the cited numbers checked, contradictions not adjudicated, code not opened, negative space not bounded). This produces **B-grade** output.
Across ~240 reviews from 2024-2026, A-grade reviews share one structural discriminator:
> **Removing the new taxonomy from an A-grade survey leaves a contribution. > Removing it from a B-grade survey leaves nothing.**
Empirical A-rate by domain (with our wave-1 + wave-2 evidence base):
| Domain | A-rate | Modal A-pattern | |---|---|---| | Math (Acta Numerica / Bull AMS / SIAM Review / Probab Surv) | ~86% | Re-derivation in unified notation; new short proofs | | Economics (JEL / Annu Rev Econ / Handbook) | ~80% | Author re-estimation on harmonized data; "stylized-fact tables" | | Engineering (Annu Rev Control/BME, PECS, ARHT) | ~73% | Author re-simulation; harmonized device spec sheets | | Physics (RMP / Living Reviews / Annu Rev Cond Matt) | ~70% | Re-derivation + cross-paper number table; per-edition updates | | Chemistry/materials (Chem Rev / Chem Soc Rev / Annu Rev Phys Chem) | ~60% | Cross-paper benchmark table; Tutorial Review structured-closing | | Earth/environment (Rev Geophys / Annu Rev Earth Planet Sci / NRE&E) | ~40% | Narrative-with-embedded-re-analysis of observational data | | CS/ML/AI surveys (arXiv survey papers) | **~13%** | Bounded corpus + author benchmarks (BetterBench template) | | Biology narrative (Nature Reviews / Annu Rev Bio / Cell / Trends) | **~7%** | Almost never — venue norm is conceptual synthesis |
Cochrane / BMJ / Lancet SRs are 100% PRISMA-compliant by editorial policy but item-level adherence is asymmetric: ~75% of Cochrane abstracts use GRADE, but only ~7.5% of nominally compliant SRs across journals do full certainty + reporting-bias assessment. **The PRISMA label is not the substance** — verify item-by-item.
Two key empirical insights from the corpus:
1. **A-grade is topic-determined, not author-determined.** Surveys of *open artifacts* (open-source models, public conference proceedings, public datasets) admit A-grade execution. Surveys of *capabilities reported by closed systems* (RLHF/alignment, frontier-model agents, healthcare LLMs, industry-disclosed tools like Aletheia) are structurally trapped at B because the survey author cannot independently re-execute cited results.
2. **Disagreement-handling is a near-universal blind spot.** 0/31 CS surveys, ~12/30 biology reviews and ~9/30 physics reviews fence-sit on contradictions. Even A-grade work routinely fails this dimension. **It is the cleanest novelty axis the agent can exploit.**
When to use this skill
Trigger when RESEARCH.md uses: *survey, review, overview, landscape, state of the art, comparative analysis, taxonomy, benchmark of benchmarks, perspective.*
Skip for primary-research projects (single experiment + paper) — those use the standard experiment / experiment_reviewer pattern directly.
Step 1 — Pick the review type explicitly
Default-narrative is the modal mistake. An autonomous agent has no editorial-gatekeeping defense, so it inherits all narrative-review failure modes (cherry-picking, confirmation bias, irreproducibility) without the defenses. **Default to PRISMA-ScR-grade documentation at minimum.**
Choose one and commit it in `notes/scope.md` before any literature load:
| Type | When to choose | Required protocol | |---|---|---| | **Audit / benchmark survey** | Field has many primary systems with reported numbers; Q is "do the claims hold?" Most CS/ML/AI SOTA survey work falls here. | BetterBench-style: bounded N, criteria list, ≥2 raters, *count* don't gesture (Reuel/Balloccu template) | | **Scoping review** | Map breadth of a heterogeneous emerging field; decide whether full SR is warranted | PRISMA-ScR (Tricco 2018), 20 items, 5-stage Arksey-O'Malley. **No quality appraisal of included sources.** | | **Systematic review** | Bounded answerable question, evidence is appraisable | PRISMA 2020 (27 items) + RoB 2 / ROBINS-I + GRADE + PROSPERO registration | | **Umbrella review** | Synthesize multiple existing SRs on a related question | AMSTAR 2 for included reviews + handle SR overlap | | **Critical narrative review** | Domain conceptual synthesis where adjudication matters more than coverage (RMP-style theoretical recap; Annu Rev Phys Chem) | Greenhalgh: explicit interpreter positioning + explicit selection logic + explicit acknowledgement of evidence not selected. **No paper-trust.** | | **Narrative-with-embedded-re-analysis** | Earth/environment / climate where review value-add is reprocessing observational datasets | `notes/datasets.md` provenance + reproducible reprocessing pipeline | | **Theoretical-unification survey** | Math/theoretical review where unifying object is the contribution (Acta Numerica template) | Re-derivation in unified notation; new short proofs of known results; compet
Read more
name: survey-methodology description: Plan-time methodology contract for survey/review/report projects. Forces an audit-grade survey instead of a paper-trust summary. Distilled from ~240 reviews (2024-2026) across 9 domain clusters — physics/RMP/Living Reviews, chemistry/materials, biology/medicine narrative + Cochrane SRs, CS/ML/AI, math/Acta Numerica, earth/environment, economics/JEL, engineering/Annual Reviews, plus PRISMA/GRADE/Cochrane protocol literature. Empirical A-grade rate by domain ranges 7% (biology narrative) to 86% (math/Acta Numerica); the discriminator is structural, not stylistic. **Read this BEFORE writing notes/plan.md for any survey-style RESEARCH.md.** compatibility: Pure prompt skill. References existing tools (spawn, escalate_authority_bound, experiment_reviewer, compile_latex). allowed-tools: Read, Edit, Write, Glob, Grep
Survey Methodology Skill
The default failure mode of an autonomous-agent survey is **paper-trust**: read N papers, organize claims into a taxonomy, ship a prose digest. The output passes type-check (it looks like a survey) but fails verification (none of the cited numbers checked, contradictions not adjudicated, code not opened, negative space not bounded). This produces **B-grade** output.
Across ~240 reviews from 2024-2026, A-grade reviews share one structural discriminator:
> **Removing the new taxonomy from an A-grade survey leaves a contribution. > Removing it from a B-grade survey leaves nothing.**
Empirical A-rate by domain (with our wave-1 + wave-2 evidence base):
| Domain | A-rate | Modal A-pattern | |---|---|---| | Math (Acta Numerica / Bull AMS / SIAM Review / Probab Surv) | ~86% | Re-derivation in unified notation; new short proofs | | Economics (JEL / Annu Rev Econ / Handbook) | ~80% | Author re-estimation on harmonized data; "stylized-fact tables" | | Engineering (Annu Rev Control/BME, PECS, ARHT) | ~73% | Author re-simulation; harmonized device spec sheets | | Physics (RMP / Living Reviews / Annu Rev Cond Matt) | ~70% | Re-derivation + cross-paper number table; per-edition updates | | Chemistry/materials (Chem Rev / Chem Soc Rev / Annu Rev Phys Chem) | ~60% | Cross-paper benchmark table; Tutorial Review structured-closing | | Earth/environment (Rev Geophys / Annu Rev Earth Planet Sci / NRE&E) | ~40% | Narrative-with-embedded-re-analysis of observational data | | CS/ML/AI surveys (arXiv survey papers) | **~13%** | Bounded corpus + author benchmarks (BetterBench template) | | Biology narrative (Nature Reviews / Annu Rev Bio / Cell / Trends) | **~7%** | Almost never — venue norm is conceptual synthesis |
Cochrane / BMJ / Lancet SRs are 100% PRISMA-compliant by editorial policy but item-level adherence is asymmetric: ~75% of Cochrane abstracts use GRADE, but only ~7.5% of nominally compliant SRs across journals do full certainty + reporting-bias assessment. **The PRISMA label is not the substance** — verify item-by-item.
Two key empirical insights from the corpus:
1. **A-grade is topic-determined, not author-determined.** Surveys of *open artifacts* (open-source models, public conference proceedings, public datasets) admit A-grade execution. Surveys of *capabilities reported by closed systems* (RLHF/alignment, frontier-model agents, healthcare LLMs, industry-disclosed tools like Aletheia) are structurally trapped at B because the survey author cannot independently re-execute cited results.
2. **Disagreement-handling is a near-universal blind spot.** 0/31 CS surveys, ~12/30 biology reviews and ~9/30 physics reviews fence-sit on contradictions. Even A-grade work routinely fails this dimension. **It is the cleanest novelty axis the agent can exploit.**
When to use this skill
Trigger when RESEARCH.md uses: *survey, review, overview, landscape, state of the art, comparative analysis, taxonomy, benchmark of benchmarks, perspective.*
Skip for primary-research projects (single experiment + paper) — those use the standard experiment / experiment_reviewer pattern directly.
Step 1 — Pick the review type explicitly
Default-narrative is the modal mistake. An autonomous agent has no editorial-gatekeeping defense, so it inherits all narrative-review failure modes (cherry-picking, confirmation bias, irreproducibility) without the defenses. **Default to PRISMA-ScR-grade documentation at minimum.**
Choose one and commit it in `notes/scope.md` before any literature load:
| Type | When to choose | Required protocol | |---|---|---| | **Audit / benchmark survey** | Field has many primary systems with reported numbers; Q is "do the claims hold?" Most CS/ML/AI SOTA survey work falls here. | BetterBench-style: bounded N, criteria list, ≥2 raters, *count* don't gesture (Reuel/Balloccu template) | | **Scoping review** | Map breadth of a heterogeneous emerging field; decide whether full SR is warranted | PRISMA-ScR (Tricco 2018), 20 items, 5-stage Arksey-O'Malley. **No quality appraisal of included sources.** | | **Systematic review** | Bounded answerable question, evidence is appraisable | PRISMA 2020 (27 items) + RoB 2 / ROBINS-I + GRADE + PROSPERO registration | | **Umbrella review** | Synthesize multiple existing SRs on a related question | AMSTAR 2 for included reviews + handle SR overlap | | **Critical narrative review** | Domain conceptual synthesis where adjudication matters more than coverage (RMP-style theoretical recap; Annu Rev Phys Chem) | Greenhalgh: explicit interpreter positioning + explicit selection logic + explicit acknowledgement of evidence not selected. **No paper-trust.** | | **Narrative-with-embedded-re-analysis** | Earth/environment / climate where review value-add is reprocessing observational datasets | `notes/datasets.md` provenance + reproducible reprocessing pipeline | | **Theoretical-unification survey** | Math/theoretical review where unifying object is the contribution (Acta Numerica template) | Re-derivation in unified notation; new short proofs of known results; compet
An autonomous research colleague — from a question to a compiled manuscript, while you sleep.
Repo: Muuuun/luxas
Other skills on luxas.
- /figure
Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly). Includes a TikZ template library covering quantum circuits (quantikz), Feynman diagrams (tikz-feynman), circuits (circuitikz), molecules (chemfig), 2D/3D plots (pgfplots), energy-level
Open skill - /matplotlib-figures
Publication-quality data plots via matplotlib with venue-specific styles. Use for generating your own figures (timelines, comparison charts, data summaries, heatmaps) that you add to the LaTeX report. Your brain prompt supplies the venue-specific style directory as
Open skill - /memory
Cross-project research memory. Deep-dive past projects' notes, record corrections, and save cross-project insights across all Luxas research projects.
Open skill - /narrative
Narrative discipline for research reports — article-type templates (empirical / feasibility / comparison / policy-zh) and the feedback revision protocol that keeps outline, prose, and figures coherent across many feedback rounds.
Open skill - /paper-figures
Extract figures from downloaded papers and include them in survey/review reports. Use when the report covers other groups' work and benefits from their architecture diagrams, experimental plots, or system schematics. Your brain prompt supplies the absolute path to the
Open skill - /qec-construct
Verifier-in-the-loop CONSTRUCTION of quantum error-correcting codes with transversal non-Clifford gates (CCZ/T). Use when the task is to INVENT or improve a code/construction (not merely search a known family). Provides a sound, cheap algebraic verifier (qverify, "QEC's Lean")
Open skill

