aaai-artifact-evaluati…
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without…
Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than
$ npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill aistats-experiments --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/aistats-experimentsContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than
name: aistats-experiments description: Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than chase leaderboards.
Use this before submission when the empirical or simulation story is not yet locked.
show practical relevance.
intervals, paired tests, or bootstrap intervals when appropriate.
settings, selection criteria, random seeds, hardware, software versions, and runtime.
theoretical assumptions and empirical setup.
simulation confirming a predicted rate outweighs five extra benchmark datasets.
they are deliberately violated, and a real-data study showing practical behavior.
dimension, noise level — matches the asymptotic regime of the theorems. A bound proven as n grows but tested only at n = 500 invites the question of relevance.
| Theoretical claim | Matching experiment | Reject pattern avoided | |---|---|---| | Convergence rate in n | Log-log error versus n with fitted slope | "Rates asserted but never plotted" | | Confidence-interval coverage | Empirical coverage across many replications | "Nominal 95 percent never verified" | | Regret bound | Cumulative regret versus horizon, with the bound curve overlaid | "Bound and trajectory never compared" | | Robustness to misspecification | Violation-severity sweep | "Guarantees hold under assumptions the experiments quietly break" |
Suppose the paper proves finite-sample type-I error control under a boundedness assumption. The matching plan: simulate under the null at several sample sizes to verify size, sweep dependence strength for power curves, then inject heavy-tailed noise that breaks boundedness to map degradation — every panel tied to a numbered theorem or remark.
are standard errors, confidence intervals, or quantiles.
[Experiment readiness] strong / adequate / weak [Claim -> evidence map] <claim: table/figure/simulation> [Missing statistical evidence] <uncertainty/test/seed/baseline> [Reproducibility gaps] <hyperparameters/compute/data/code> [Decision-critical next run] <one experiment or simulation>
Stanford REAP × CoPaper.AI · 由斯坦福实证方法论团队精选与维护 访问 copaper.ai 微信:CoPaper.AI 按 11 个主流学科板块覆盖 经管与商科 社会科学 人文学科 数学与物理科学 生命科学 医学与健康 工程与技术 计算机科学与 AI 体育科学 点击任一学科名可跳转到对应说明;每类下的代表子领域在正文总览中完整列出。下方封面墙按 venue 导航,完整分类见覆盖一览。 🧭 布局指南 · 📚 Skill Pack 一览 · ⚡ 如何使用 · 🧪 自动实证
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without…
Use when drafting an AAAI author response (rebuttal) under the single short character-limited author-feedback window, the no-URL rule, no-new-results guidance,…
Use when preparing an accepted AAAI paper for camera-ready source submission to AAAI Press, including proceedings page limits, two-column template compliance,…
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human…
Use when positioning an AAAI paper's novelty against archival work, contemporaneous arXiv or workshop papers, and AAAI/IJCAI/NeurIPS/ICML/ICLR neighbors across…
Use when strengthening an AAAI paper's reproducibility checklist (placed after references), experimental traceability, seed and hyperparameter reporting,…