pipeline
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
End-to-end data analysis workflow in R or Python — from exploration through regression to publication-ready tables and figures. Make sure to use this skill whenever the user wants to run any empirical analysis, write analysis code, or produce output from data. Triggers include:
$ npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill data-analysis --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/data-analysisContext preview
The summary Claude sees to decide when to auto-load this skill.
End-to-end data analysis workflow in R or Python — from exploration through regression to publication-ready tables and figures. Make sure to use this skill whenever the user wants to run any empirical analysis, write analysis code, or produce output from data. Triggers include:
name: data-analysis description: >- End-to-end data analysis workflow in R or Python — from exploration through regression to publication-ready tables and figures. Make sure to use this skill whenever the user wants to run any empirical analysis, write analysis code, or produce output from data. Triggers include: "analyze this data", "run a regression", "write R code for this", "write Python code for this", "I have a dataset", "help me with this regression", "run a DiD", "run an RDD", "event study", "IV regression", "fit a model", "produce a table", "make a figure", "explore my data", or any request involving a dataset path or empirical estimation. argument-hint: "[dataset path or description of analysis goal]" allowed-tools: ["Read", "Grep", "Glob", "Write", "Edit", "Bash", "Task", "AskUserQuestion"]
Run an end-to-end data analysis in R or Python: load, explore, analyze, and produce publication-ready output.
**Input:** `$ARGUMENTS` — a dataset path (e.g., `data/county_panel.csv`) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").
---
Determine language from `$ARGUMENTS` or ask the user:
---
1. Create R script with proper header (title, author, purpose, inputs, outputs) 2. Load required packages at top (`library()`, never `require()`) 3. Set seed once at top: `set.seed(42)` 4. Create output directories: `dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)` 5. Load and inspect the dataset
**Tables:** `modelsummary` (preferred) or `stargazer` — export `.tex` and `.html` **Figures:** `ggplot2` with project theme; explicit `ggsave(width = X, height = Y)`; save as `.pdf` and `.png`; add `bg = "transparent"` only if output is for Beamer slides
1. `saveRDS()` for all key objects 2. Run the `r-reviewer` agent: *"Review the script at scripts/R/[script_name].R"* 3. Address Critical and High issues before presenting results
# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================
# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)
set.seed(42)
dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)
# 1. Data Loading ----
# 2. Exploratory Analysis ----
# 3. Main Analysis ----
# 4. Tables and Figures ----
# 5. Export -------
1. Create Python script with header (title, author, purpose, inputs, outputs) 2. Import all packages at the top of the file 3. Set seeds: `np.random.seed(42)` and `random.seed(42)` 4. Create output directories: `Path("output/analysis").mkdir(parents=True, exist_ok=True)` 5. Load and inspect the dataset with `pandas`
**Tables:** Format with `pandas` and export via `.to_latex()` or `stargazer` (Python port) **Figures:** `matplotlib`/`seaborn`; explicit `fig.savefig(path, dpi=300, bbox_inches="tight")`; save as `.pdf` and `.png`
1. `joblib.dump(model, "output/model.pkl")` for fitted models 2. `df_results.to_parquet("output/results.parquet")` for DataFrames 3. Review the script manually against the Python checklist bel
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。 English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位) 🌐 语言: English |
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Use when the user asks to run a full empirical / causal analysis in Python — by default in the style of an applied economics paper (AER / QJE / JPE / ReStud /…
Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest +…
Classical end-to-end empirical analysis workflow in the traditional Stata ecosystem — native Stata + reghdfe + ivreg2 + csdid + did_imputation +…
Classical end-to-end empirical analysis workflow in the modern tidyverse + econometrics R ecosystem — dplyr + tidyr + haven + fixest + sandwich + lmtest +…
Systematic writing framework for philosophy and interdisciplinary academic papers from optimized outline to submission-ready manuscript. Use when users want…