pipeline
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
This skill covers causal machine learning methods in applied economics and quantitative social science. Use when implementing or choosing between modern ML-based causal estimators — including double machine learning, DML, partially linear models, interactive regression models,
$ npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill causal-ml --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/causal-mlContext preview
The summary Claude sees to decide when to auto-load this skill.
This skill covers causal machine learning methods in applied economics and quantitative social science. Use when implementing or choosing between modern ML-based causal estimators — including double machine learning, DML, partially linear models, interactive regression models,
name: causal-ml argument-hint: "<estimator or method choice>" description: >- This skill covers causal machine learning methods in applied economics and quantitative social science. Use when implementing or choosing between modern ML-based causal estimators — including double machine learning, DML, partially linear models, interactive regression models, cross-fitting, Neyman orthogonality, debiased ML, causal forests, generalized random forest, GRF, honest causal trees, AIPW with machine learning, doubly robust with machine learning, DR-Learner, T-Learner, S-Learner, X-Learner, meta-learners, heterogeneous treatment effects, conditional average treatment effect, CATE, HTE, high-dimensional controls, LASSO controls, post-LASSO, post-double selection, Belloni-Chernozhukov-Hansen, Riesz representer, Chernozhukov, sample splitting, econml, DoubleML package, or any combination of machine learning and causal inference.
Reference for semiparametric ML estimators: DML with cross-fitting, generalized random forests, debiased regularization, and nuisance function approximation. Covers Neyman-orthogonal moment conditions, sample splitting, plug-in bias correction, and heterogeneous treatment effects.
Use when the user is:
Skip when:
---
| Dimension | Traditional (IV, DiD, RDD) | Causal ML | |-----------|--------------------------|-----------| | Functional form | Parametric | Nonparametric / semi-parametric | | High-dimensional controls | Problematic | Native support | | Heterogeneous effects | Secondary (subgroup analysis) | Primary estimand (CATE) | | Sample requirements | Moderate N | ML nuisance needs large N | | Identification | Explicit (IV, DiD, RCT) | Same assumptions — ML is estimation, not identification |
**Critical point:** Causal ML does not relax identification assumptions. If you need a valid instrument, parallel trends, or no unmeasured confounding, those must still hold.
---
DML (Chernozhukov et al. 2018) fixes regularization bias in naive ML-in-regression. Partial out controls X from both Y and D using separate ML nuisance models, then regress residuals. Two properties: **Neyman orthogonality** (moment condition locally insensitive to nuisance error) and **cross-fitting** (prevents overfitting bias).
**PLR** (Partially Linear Regression): $Y = \theta D + g(X) + \varepsilon$. Workhorse for continuous or binary D with ATE under selection on observables. **IRM** (Interactive Regression Model): relaxes additive separability for binary D with heterogeneous effects.
Full implementation (Python/R code, cross-fitting from scratch, diagnostics) in `references/dml.md`.
Causal forests (Wager-Athey 2018; Athey-Tibshirani-Wager 2019) estimate CATE $\tau(x) = E[Y(1)-Y(0)|X=x]$ using **honest** forests (structure learned on one subsample, effects estimated on another). Use when CATE is the primary estimand and n $\geq$ 2,000. Always run the calibration test before reporting heterogeneity.
R (`grf`) and Python (`econml`) implementations, ATE/ATT extraction, BLP projections in `references/grf-meta-learners.md`.
Decompose CATE estimation into supervised learning sub-problems. **DR-Learner** (Kennedy 2023): best properties when both nuisance models are well-specified. **T-Learner**: simplest baseline. **X-Learner**: designed for imbalanced treatment. For applied work: DR-Learner primary, T-Learner benchmark. Large disagreement signals nuisance model problems.
All implementations in `references/grf-meta-learners.md`.
**PDS-LASSO** (Belloni-Chernozhukov-Hansen 2014): separate LASSOes of Y on X and D on X, union of selected variables, then OLS. Works at moderate n (~200 with sparse confounders). See `references/high-dim-cross-fitting.md`.
Before reporting CATE, test for genuine heterogeneity using BLP calibration test. Do not report heterogeneous effects if calibration test fails (p > 0.10). See `references/hte-inference.md`.
---
1. n < 500? → Use standard methods (causal-inference skill) 2. High-dim controls (p > 20), want ATE? → PDS-LASSO or DML-PLR; binary D → DML-IRM 3. CATE is primary estimand? → Causal Forest (large n) or DR-Learner (doubly robust) 4. Endogenous treatment with instrument? → DML-PLIV 5. Treatment is rare/imbalanced? → X-Learner 6. Quick benchmark? → Always compute T-Learner as baseline
| Method | Estimand | Python | R | Min n | Key diagnostic | |--------|----------|--------|---|-------|----------------| | DML-PLR | ATE | `doubleml`, `econml` | `DoubleML` | ~500 | Nuisance R², res
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。 English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位) 🌐 语言: English |
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Use when the user asks to run a full empirical / causal analysis in Python — by default in the style of an applied economics paper (AER / QJE / JPE / ReStud /…
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Classical end-to-end empirical analysis workflow in the traditional Stata ecosystem — native Stata + reghdfe + ivreg2 + csdid + did_imputation +…
Classical end-to-end empirical analysis workflow in the modern tidyverse + econometrics R ecosystem — dplyr + tidyr + haven + fixest + sandwich + lmtest +…
Systematic writing framework for philosophy and interdisciplinary academic papers from optimized outline to submission-ready manuscript. Use when users want…