pipeline
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Agent D4 - Measurement Instrument Developer - Scale construction and psychometric validation. Covers item development, validity evidence, and reliability testing for social science research.
$ npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill d4 --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/d4Context preview
The summary Claude sees to decide when to auto-load this skill.
Agent D4 - Measurement Instrument Developer - Scale construction and psychometric validation. Covers item development, validity evidence, and reliability testing for social science research.
name: d4 description: | Agent D4 - Measurement Instrument Developer - Scale construction and psychometric validation. Covers item development, validity evidence, and reliability testing for social science research. version: "12.0.1"
`diverga_check_prerequisites("d4")` → must return `approved: true` If not approved → AskUserQuestion for each missing checkpoint (see `.claude/references/checkpoint-templates.md`)
Read `.research/decision-log.yaml` directly to verify prerequisites. Conversation history is last resort.
---
Develop psychometrically sound measurement instruments (scales, questionnaires, surveys) for social science research, ensuring construct validity, reliability, and appropriate psychometric properties.
item_writing_guidelines:
clarity:
- Use simple, direct language
- One idea per item
- Appropriate reading level for target population
- Avoid jargon and technical terms
neutrality:
- Avoid leading questions
- No double-barreled questions
- No loaded language
- Balanced positive and negative items
specificity:
- Concrete behaviors over abstract traits
- Specific time frames when relevant
- Clear referent (who/what is being asked about)
avoid:
- Double negatives ("not unlikely")
- Ambiguous frequency words ("sometimes", "often")
- Extreme modifiers ("always", "never")
- Hypothetical scenarios without contextresponse_formats:
likert_scale:
points: "5-7 points (odd for neutral option)"
labels:
all_points: "More precise but takes more space"
endpoints_only: "Cleaner but assumes equal intervals"
direction: "Maintain consistency across entire scale"
examples:
agreement: ["Strongly Disagree", "Disagree", "Neutral", "Agree", "Strongly Agree"]
frequency: ["Never", "Rarely", "Sometimes", "Often", "Always"]
semantic_differential:
format: "Bipolar adjective pairs with 7-point scale"
spacing: "Equal visual spacing between points"
example: |
Good ___:___:___:___:___:___:___ Bad
1 2 3 4 5 6 7
visual_analog:
format: "Continuous line with endpoints labeled"
scoring: "Convert to 0-100 scale"
use_case: "Pain, mood, subjective states"
forced_choice:
format: "Choose between two statements"
use_case: "Reduce social desirability bias"
example: "Which describes you better? A or B"
ranking:
format: "Order items by importance/preference"
limitation: "Complex for respondents, limited items (max 7-10)"
checklist:
format: "Select all that apply"
use_case: "Behaviors, experiences, symptoms"scale_development_stages:
stage_1_conceptualization:
duration: "2-4 weeks"
activities:
- Define construct with theoretical grounding
- Literature review for existing scales
- Develop operational definition
- Identify dimensions/facets if multidimensional
outputs:
- Construct definition document
- Conceptual framework diagram
- Decision: Unidimensional vs. Multidimensional
stage_2_item_generation:
duration: "3-6 weeks"
guidelines:
initial_pool_size: "3-4x the target final number of items"
sources:
- Literature review (adapt existing items)
- Expert interviews (domain specialists)
- Target population interviews (actual language used)
- Theory (deductive approach)
coverage:
- Ensure all facets/dimensions represented
- Balance positive and negative items
- Range from low to high levels of construct
outputs:
- Item pool (40-60 items for 10-15 final items)
- Item classification by dimension
stage_3_expert_review:
duration: "2-3 weeks"
participants: "5-10 content experts"
method:
content_validity_ratio:
formula: "CVR = (ne - N/2) / (N/2)"
interpretation: "CVR > 0.62 for N=7 experts (Lawshe, 1975)"
decision_rule: "Retain items with CVR above threshold"
expert_ratings:
- Relevance (1=Not relevant, 4=Highly relevant)
- Clarity (1=Not clear, 4=Very clear)
- Representativeness of dimension
qualitative_feedback:
- Wording suggestions
- Missing content areas
- Cultural appropriateness
outputs:
- Revised item pool (30-40 items)
- Content validity evidence
stage_4_cognitive_interview:
duration: "2-3 weeks"
participants: "8-15 members of target population"
method:
think_aloud:
- Read item aloud
- Say what they're thinking
- Explain their answer choice
verbal_probing:
comprehension: "What does this question mean to you?"
retrieval: "How did you arrive at your answer?"
judgment: "How easy or difficult was it to answer?"
response: "Why did you choose that option?"
analysis:
- Identify comprehension problems
- Detect unintended interpretations
- Find offensive/inappropriate items
outputs:
- Revised items based on feedback
- Response process validity evidence
stage_5_pilot_test:
duration: "4-6 weeks"
sample_size: "150-300 (5-10 per item minimum)"
recruitment: "Representative of target population"
analysis:
descriptive_statistics:
- Mean, SD, skewness, kurtosis for each item
- Floor/ceiling effects (>15% at extremes = problem)
- Missing data patterns
item_analysis:
item_total_correlation:
threshold: "r > .📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。 English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位) 🌐 语言: English |
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Use when the user asks to run a full empirical / causal analysis in Python — by default in the style of an applied economics paper (AER / QJE / JPE / ReStud /…
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Classical end-to-end empirical analysis workflow in the traditional Stata ecosystem — native Stata + reghdfe + ivreg2 + csdid + did_imputation +…
Classical end-to-end empirical analysis workflow in the modern tidyverse + econometrics R ecosystem — dplyr + tidyr + haven + fixest + sandwich + lmtest +…
Systematic writing framework for philosophy and interdisciplinary academic papers from optimized outline to submission-ready manuscript. Use when users want…