aaai-artifact-evaluati…
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without…
Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data
$ npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/acl-artifact-evaluationContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data
name: acl-artifact-evaluation description: Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-acceptance public release.
Use this to plan the evidence package around an ACL paper. ACL has no separate artifact-badge track; instead, artifact scrutiny is folded into review through the supplement archive and Section B ("scientific artifacts") of the Responsible NLP checklist, which reviewers cross-check against the PDF.
suites, adversarial sets.
cheapest way to make an LLM paper checkable without GPUs.
consent text, compensation description.
tracked cloud storage are not acceptable, and any linked page must be anonymous.
notebook author fields, license headers, dataset hosting pages, README contact lines.
must stand alone; the archive is for verification, not for essential content.
| Responsible NLP item (Section B) | Artifact implication | |---|---| | Cited creators + versions of used artifacts | Pin dataset/model versions in the README and bibliography | | License / terms of use stated | Include the license you release under *and* those you consumed under | | Use consistent with intended use | Justify research use of scraped or user-generated data | | PII and offensive content handled | Describe scanning/anonymization steps actually performed | | Documentation of domains, languages, demographics | Ship a data statement or datasheet, not just row counts | | Statistics on splits reported | Train/dev/test sizes in both paper and README |
Checklist answers contradicted by the archive read as misleading information — grounds for desk rejection under ARR policy, and a credibility wound even when not enforced.
1. The README — it has roughly one minute to orient them. 2. Prompt files and evaluation scripts, for any LLM claim: exact prompts, decoding parameters, and scoring code are the reproduction spine. 3. Annotation guidelines, for any dataset or human-eval claim: reviewers judge whether the labels could possibly mean what the paper says. 4. A sample of the data itself — quality problems visible in twenty random examples have sunk otherwise strong resource papers.
A hypothetical paper releases a 7-language reading-comprehension test suite built from news text plus a baseline evaluation of five LLMs.
filtering pipeline as runnable code, since "web text" alone fails checklist item B on documentation.
statistics; multilingual annotation quality is the first attack surface.
re-score without API keys.
can be audited later.
anonymous supplement -> public repo + dataset page -> archived, versioned release (review-time) (camera-ready links) (DOI/hub artifact, cited version)
Post-acceptance, register the artifact where your community actually looks (model/dataset hubs, a maintained repo), state the license explicitly, and put the citation-of-record (the Anthology entry) in the README.
Run these before zipping, on a copy:
# authorship trails in code and docs grep -ri "yourname\|yourlab\|university" . --include="*.py" --include="*.md" # git history and remotes leak owners rm -rf .git; # or re-init a fresh repo for the archive copy # notebook metadata carries usernames and kernel paths jupyter nbconvert --clear-output --inplace *.ipynb # absolute paths in configs and logs grep -r "/home/\|/Users/" . | head
Then check the parts tools miss: license headers naming the lab, dataset hosting pages with institutional branding, model cards listing maintainers, and README badges pointing at owner-named CI.
cached datasets, and virtualenvs; describe big assets and provide them at camera-ready instead.
result — reviewers grant roughly a minute before giving up.
OpenReview upload limits and accepted fields vary by cycle, so check the live form rather than last cycle's.
[Artifact role] anonymous supplement / camera-ready release / public benchmark [Contents] <code/data/prompts/outputs/guidelines> [Checklist alignment] <Section B items satisfied vs missing> [Anonymity findings] <paths/metadata/hosting leaks> [Release plan] <post-acceptance registry, license, versioning>
Stanford REAP × CoPaper.AI · 由斯坦福实证方法论团队精选与维护 访问 copaper.ai 微信:CoPaper.AI 按 11 个主流学科板块覆盖 经管与商科 社会科学 人文学科 数学与物理科学 生命科学 医学与健康 工程与技术 计算机科学与 AI 体育科学 点击任一学科名可跳转到对应说明;每类下的代表子领域在正文总览中完整列出。下方封面墙按 venue 导航,完整分类见覆盖一览。 🧭 布局指南 · 📚 Skill Pack 一览 · ⚡ 如何使用 · 🧪 自动实证
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without…
Use when drafting an AAAI author response (rebuttal) under the single short character-limited author-feedback window, the no-URL rule, no-new-results guidance,…
Use when preparing an accepted AAAI paper for camera-ready source submission to AAAI Press, including proceedings page limits, two-column template compliance,…
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human…
Use when positioning an AAAI paper's novelty against archival work, contemporaneous arXiv or workshop papers, and AAAI/IJCAI/NeurIPS/ICML/ICLR neighbors across…
Use when strengthening an AAAI paper's reproducibility checklist (placed after references), experimental traceability, seed and hyperparameter reporting,…