/b2
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias,
$ npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill b2 --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/b2
Context preview
The summary Claude sees to decide when to auto-load this skill.
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias,
SKILL.md
b2.SKILL.mdname: b2
description: |
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment
Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies
Use when: appraising study quality, assessing risk of bias, grading evidence
Triggers: quality appraisal, RoB, GRADE, Newcastle-Ottawa, risk of bias, methodological quality
version: "12.0.1"
⛔ Prerequisites (v8.2 — MCP Enforcement)
`diverga_check_prerequisites("b2")` → must return `approved: true` If not approved → AskUserQuestion for each missing checkpoint (see `.claude/references/checkpoint-templates.md`)
Checkpoints During Execution
- 🟠 CP_QUALITY_REVIEW → `diverga_mark_checkpoint("CP_QUALITY_REVIEW", decision, rationale)`
Fallback (MCP unavailable)
Read `.research/decision-log.yaml` directly to verify prerequisites. Conversation history is last resort.
---
Evidence Quality Appraiser
**Agent ID**: 06 **Category**: B - Literature & Evidence **VS Level**: Enhanced (3-Phase) **Tier**: Core **Icon**: 🔬
Overview
Systematically evaluates methodological quality and risk of bias in individual studies. Selects and applies appropriate assessment tools based on study design type.
Applies **VS-Research methodology** to go beyond mechanical tool application, providing differentiated quality evaluation strategies tailored to research context and purpose.
VS-Research 3-Phase Process (Enhanced)
Phase 1: Modal Quality Assessment Approach Identification
**Purpose**: Recognize limitations of mechanical tool application
⚠️ **Modal Warning**: The following are the most predictable quality assessment approaches:
| Modal Approach | T-Score | Limitation |
|----------------|---------|------------|
| "RCT → Apply RoB 2.0" | 0.90 | Automatic matching ignoring context |
| "Observational → Apply NOS" | 0.88 | Ignores tool limitations |
| "Report GRADE rating only" | 0.85 | Rating rationale unclear |
➡️ Tool application is baseline. Proceeding with context-adaptive assessment.
Phase 2: Context-Adaptive Evaluation Strategy
**Purpose**: Present evaluation approaches suited to research purpose and context
**Direction A** (T ≈ 0.7): Standard tool + contextual interpretation
- Standard tool application + domain-specific weighting
- Suitable for: General systematic reviews
**Direction B** (T ≈ 0.4): Multi-tool triangulation
- Simultaneous application of multiple tools + discrepancy analysis
- Additional field-specific quality criteria
- Suitable for: Methodology papers, high-quality reviews
**Direction C** (T < 0.3): Purpose-specific evaluation
- Differentiated criteria by meta-analysis purpose
- Propose new evaluation dimensions (reproducibility, transparency)
- Suitable for: Methodological innovation, guideline development
Phase 4: Recommendation Execution
Based on **selected evaluation strategy**: 1. State tool selection rationale 2. Domain-specific detailed assessment + interpretive commentary 3. Meta-analysis utilization recommendations 4. Sensitivity analysis necessity determination
---
Quality Assessment Typicality Score Reference Table
T > 0.8 (Modal - Supplementation Required):
├── Study type → Standard tool automatic matching
├── Yes/No per checklist item
├── Report only total score or rating
└── Judgment rationale unclear
T 0.5-0.8 (Established - Add Interpretation):
├── Specific rationale per domain
├── Interpret meaning in research context
├── Meta-analysis inclusion/exclusion recommendation
└── Sensitivity analysis necessity determination
T 0.3-0.5 (In-depth - Recommended):
├── Multi-tool triangulation
├── Additional field-specific criteria
├── Quality-effect size relationship analysis
└── Rating uncertainty quantification
T < 0.3 (Innovative - For Leading Research):
├── Propose new evaluation dimensions
├── Critical discussion of tool limitations
├── Purpose-specific evaluation framework
└── Quality assessment uncertainty propagation
When to Use
- Evaluating included studies in systematic reviews
- Verifying study quality before meta-analysis
- Assessing evidence for evidence-based decision making
- Judging reliability of research findings
Core Functions
1. **Study Type-Specific Tool Selection**
- RCT: Cochrane Risk of Bias 2.0
- Observational studies: Newcastle-Ottawa Scale, ROBINS-I
- Qualitative studies: CASP, JBI Critical Appraisal
- Mixed methods: MMAT
2. **Risk of Bias Assessment**
- Domain-specific bias evaluation
- Overall risk of bias judgment
- Evidence-based determination
3. **GRADE Certainty Rating**
- Certainty of evidence assessment
- Identify upgrade/downgrade factors
- Support recommendation strength judgment
4. **Quality Summary Visualization**
- Traffic light plot
- Summary of findings table
Assessment Tool Library
RCT: Cochrane Risk of Bias 2.0
| Domain | Assessment Content | |--------|-------------------| | D1 | Bias arising from randomization process | | D2 | Bias due to deviations from intended interventions | | D3 | Bias due to missing outcome data | | D4 | Bias in measurement of outcome | | D5 | Bias in selection of reported result |
**Judgment**: Low risk / Some concerns / High risk
Observational Studies: Newcastle-Ottawa Scale
| Domain | Items | Points | |--------|-------|--------| | Selection | Representativeness of exposed cohort | ★ | | | Selection of non-exposed cohort | ★ | | | Ascertainment of exposure | ★ | | | Demonstration outcome not present at start | ★ | | Comparability | Comparability of cohorts | ★★ | | Outcome | Assessment of outcome | ★ | | | Adequate follow-up length | ★ | | | Adequacy of follow-up | ★ |
**Total Score**: /9 points
Qualitative Studies: CASP Checklist
1. Was there a clear statement of aims? 2. Is a qualitative methodology appropriate? 3. Was the research design appropriate? 4. Was the recruitment strategy appropriate? 5. W
Read more
name: b2 description: | VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias, grading evidence Triggers: quality appraisal, RoB, GRADE, Newcastle-Ottawa, risk of bias, methodological quality version: "12.0.1"
⛔ Prerequisites (v8.2 — MCP Enforcement)
`diverga_check_prerequisites("b2")` → must return `approved: true` If not approved → AskUserQuestion for each missing checkpoint (see `.claude/references/checkpoint-templates.md`)
Checkpoints During Execution
- 🟠 CP_QUALITY_REVIEW → `diverga_mark_checkpoint("CP_QUALITY_REVIEW", decision, rationale)`
Fallback (MCP unavailable)
Read `.research/decision-log.yaml` directly to verify prerequisites. Conversation history is last resort.
---
Evidence Quality Appraiser
**Agent ID**: 06 **Category**: B - Literature & Evidence **VS Level**: Enhanced (3-Phase) **Tier**: Core **Icon**: 🔬
Overview
Systematically evaluates methodological quality and risk of bias in individual studies. Selects and applies appropriate assessment tools based on study design type.
Applies **VS-Research methodology** to go beyond mechanical tool application, providing differentiated quality evaluation strategies tailored to research context and purpose.
VS-Research 3-Phase Process (Enhanced)
Phase 1: Modal Quality Assessment Approach Identification
**Purpose**: Recognize limitations of mechanical tool application
⚠️ **Modal Warning**: The following are the most predictable quality assessment approaches: | Modal Approach | T-Score | Limitation | |----------------|---------|------------| | "RCT → Apply RoB 2.0" | 0.90 | Automatic matching ignoring context | | "Observational → Apply NOS" | 0.88 | Ignores tool limitations | | "Report GRADE rating only" | 0.85 | Rating rationale unclear | ➡️ Tool application is baseline. Proceeding with context-adaptive assessment.
Phase 2: Context-Adaptive Evaluation Strategy
**Purpose**: Present evaluation approaches suited to research purpose and context
**Direction A** (T ≈ 0.7): Standard tool + contextual interpretation - Standard tool application + domain-specific weighting - Suitable for: General systematic reviews **Direction B** (T ≈ 0.4): Multi-tool triangulation - Simultaneous application of multiple tools + discrepancy analysis - Additional field-specific quality criteria - Suitable for: Methodology papers, high-quality reviews **Direction C** (T < 0.3): Purpose-specific evaluation - Differentiated criteria by meta-analysis purpose - Propose new evaluation dimensions (reproducibility, transparency) - Suitable for: Methodological innovation, guideline development
Phase 4: Recommendation Execution
Based on **selected evaluation strategy**: 1. State tool selection rationale 2. Domain-specific detailed assessment + interpretive commentary 3. Meta-analysis utilization recommendations 4. Sensitivity analysis necessity determination
---
Quality Assessment Typicality Score Reference Table
T > 0.8 (Modal - Supplementation Required): ├── Study type → Standard tool automatic matching ├── Yes/No per checklist item ├── Report only total score or rating └── Judgment rationale unclear T 0.5-0.8 (Established - Add Interpretation): ├── Specific rationale per domain ├── Interpret meaning in research context ├── Meta-analysis inclusion/exclusion recommendation └── Sensitivity analysis necessity determination T 0.3-0.5 (In-depth - Recommended): ├── Multi-tool triangulation ├── Additional field-specific criteria ├── Quality-effect size relationship analysis └── Rating uncertainty quantification T < 0.3 (Innovative - For Leading Research): ├── Propose new evaluation dimensions ├── Critical discussion of tool limitations ├── Purpose-specific evaluation framework └── Quality assessment uncertainty propagation
When to Use
- Evaluating included studies in systematic reviews
- Verifying study quality before meta-analysis
- Assessing evidence for evidence-based decision making
- Judging reliability of research findings
Core Functions
1. **Study Type-Specific Tool Selection**
- RCT: Cochrane Risk of Bias 2.0
- Observational studies: Newcastle-Ottawa Scale, ROBINS-I
- Qualitative studies: CASP, JBI Critical Appraisal
- Mixed methods: MMAT
2. **Risk of Bias Assessment**
- Domain-specific bias evaluation
- Overall risk of bias judgment
- Evidence-based determination
3. **GRADE Certainty Rating**
- Certainty of evidence assessment
- Identify upgrade/downgrade factors
- Support recommendation strength judgment
4. **Quality Summary Visualization**
- Traffic light plot
- Summary of findings table
Assessment Tool Library
RCT: Cochrane Risk of Bias 2.0
| Domain | Assessment Content | |--------|-------------------| | D1 | Bias arising from randomization process | | D2 | Bias due to deviations from intended interventions | | D3 | Bias due to missing outcome data | | D4 | Bias in measurement of outcome | | D5 | Bias in selection of reported result |
**Judgment**: Low risk / Some concerns / High risk
Observational Studies: Newcastle-Ottawa Scale
| Domain | Items | Points | |--------|-------|--------| | Selection | Representativeness of exposed cohort | ★ | | | Selection of non-exposed cohort | ★ | | | Ascertainment of exposure | ★ | | | Demonstration outcome not present at start | ★ | | Comparability | Comparability of cohorts | ★★ | | Outcome | Assessment of outcome | ★ | | | Adequate follow-up length | ★ | | | Adequacy of follow-up | ★ |
**Total Score**: /9 points
Qualitative Studies: CASP Checklist
1. Was there a clear statement of aims? 2. Is a qualitative methodology appropriate? 3. Was the research design appropriate? 4. Was the recruitment strategy appropriate? 5. W
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。 English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位) 🌐 语言: English |
Other skills on auto-empirical-research-skills.
- /pipeline
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest + rdrobust + econml + causalml + matplotlib/seaborn. **Defaults to economics empirical-paper style** (AER / QJE / AEJ) —
Open skill - /pipeline
Classical end-to-end empirical analysis workflow in the modern tidyverse + econometrics R ecosystem — dplyr + tidyr + haven + fixest + sandwich + lmtest + clubSandwich + AER + ivreg + did + bacondecomp + HonestDiD + eventstudyr + rdrobust + rddensity + Synth + gsynth + synthdid
Open skill - /pipeline
Classical end-to-end empirical analysis workflow in the traditional Stata ecosystem — native Stata + reghdfe + ivreg2 + csdid + did_imputation + eventstudyinteract + sdid + rdrobust + rddensity + synth + synth_runner + psmatch2 + teffects + ebalance + coefplot + esttab + asdoc +
Open skill - /00-Full-empirical-analysis-skill_StatsPAI
Use when the user asks to run a full empirical / causal analysis in Python — by default in the style of an applied economics paper (AER / QJE / JPE / ReStud / AEJ) with DID / RD / IV / SCM / DML / matching, written-out estimating equation + identifying assumption, Table 1 /
Open skill - /00.1-Full-empirical-analysis-skill_Python
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest + rdrobust + econml + causalml + matplotlib/seaborn. **Defaults to economics empirical-paper style** (AER / QJE / AEJ) —
Open skill - /00.2-Full-empirical-analysis-skill_Stata
Classical end-to-end empirical analysis workflow in the traditional Stata ecosystem — native Stata + reghdfe + ivreg2 + csdid + did_imputation + eventstudyinteract + sdid + rdrobust + rddensity + synth + synth_runner + psmatch2 + teffects + ebalance + coefplot + esttab + asdoc +
Open skill

