/aamas-artifact-evaluation
Use when packaging AAMAS multiagent code, environments, opponent and population sets, random seeds, game definitions, and logs as anonymous supplementary evidence or a public post-acceptance release, even without a separate artifact badge, so that game-theory and MARL reviewers
$ npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill aamas-artifact-evaluation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/aamas-artifact-evaluation
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when packaging AAMAS multiagent code, environments, opponent and population sets, random seeds, game definitions, and logs as anonymous supplementary evidence or a public post-acceptance release, even without a separate artifact badge, so that game-theory and MARL reviewers
SKILL.md
aamas-artifact-evaluation.SKILL.mdname: aamas-artifact-evaluation
description: Use when packaging AAMAS multiagent code, environments, opponent and population sets, random seeds, game definitions, and logs as anonymous supplementary evidence or a public post-acceptance release, even without a separate artifact badge, so that game-theory and MARL reviewers can inspect and re-run the interaction claims.
AAMAS Artifact Evaluation
Use this for evidence packaging around AAMAS. Because the venue is about interaction, an artifact must make a *multiagent* claim inspectable: the game, the other agents, and the protocol, not just a single trained model.
Artifact plan
- Decide what a reviewer needs to believe the interaction claim: game or environment code,
opponent/population definitions, the training regime, seeds, payoff logs, proofs, or qualitative episode traces.
- Keep decision-critical evidence in the main paper or appendix; optional bulk runs can live in
the supplementary zip.
- Anonymize repository history, paths, environment names, license headers, cluster paths, and
commit authors for the review version.
- Include a minimal reproduction map: environment build, dependencies, hardware, commands,
expected outputs, per-run wall-clock, seeds, and known nondeterminism (especially in self-play).
- For a deployed or human-subject setting, give enough provenance for credible reproduction
without violating data-use terms.
- After acceptance, replace anonymous archives with a public, licensed, citable artifact.
What AAMAS evidence reviewers open first
The single fact that shapes packaging: a reviewer will re-run a small **game** far sooner than they will retrain a large policy, so make the strategic core turnkey before polishing anything.
| Claim type | First artifact inspected | Common failure caught | |---|---|---| | Convergence to an equilibrium | The game definition and the learning-rule code | Solution concept named in the paper but not encoded in the evaluation | | Emergent cooperation/defection | The environment and reward specification | Result depends on an undocumented reward-shaping constant | | Beats other agents | The opponent/population set and match protocol | Only self-play reported; no held-out opponents | | Mechanism is truthful | The payment rule plus a strategic-deviation test | No script that lets an agent try to game the mechanism |
Worked vignette: packaging a self-play study
A hypothetical submission claims a learning rule that converges to a correlated equilibrium in a repeated congestion game, shown by self-play.
- Ship the game as one parameterized generator (number of agents, capacity, payoff scale)
rather than constants buried in a notebook, so reviewers can vary the interaction.
- Record the exact seed sequence and replication count behind every convergence plot; an
equilibrium-convergence claim without seeds is unfalsifiable.
- Emit payoff and regret tables directly from logged results so PDF and artifact numbers cannot
drift.
- Include a strategic-deviation harness: a script that drops in a non-conforming agent and
measures whether it profits, because that is exactly what a game-theory reviewer will try.
Calibration anchors
- Supplement inspection at AAMAS is at reviewer discretion; assume only the README and one entry
script get opened, and design the top level accordingly.
- Supplement size and format caps vary by cycle (25 MB single zip in 2026); verify against the
current OpenReview form rather than a past year.
Output format
[Artifact role] anonymous supplement / camera-ready release / public archive
[Contents] <game/env/opponents/seeds/proofs/logs>
[Anonymity risks] <paths/metadata/licenses/URLs>
[Reproduction level] turnkey / scripted / descriptive / weak
[Fixes before upload] <ordered list>
Read more
name: aamas-artifact-evaluation description: Use when packaging AAMAS multiagent code, environments, opponent and population sets, random seeds, game definitions, and logs as anonymous supplementary evidence or a public post-acceptance release, even without a separate artifact badge, so that game-theory and MARL reviewers can inspect and re-run the interaction claims.
AAMAS Artifact Evaluation
Use this for evidence packaging around AAMAS. Because the venue is about interaction, an artifact must make a *multiagent* claim inspectable: the game, the other agents, and the protocol, not just a single trained model.
Artifact plan
- Decide what a reviewer needs to believe the interaction claim: game or environment code,
opponent/population definitions, the training regime, seeds, payoff logs, proofs, or qualitative episode traces.
- Keep decision-critical evidence in the main paper or appendix; optional bulk runs can live in
the supplementary zip.
- Anonymize repository history, paths, environment names, license headers, cluster paths, and
commit authors for the review version.
- Include a minimal reproduction map: environment build, dependencies, hardware, commands,
expected outputs, per-run wall-clock, seeds, and known nondeterminism (especially in self-play).
- For a deployed or human-subject setting, give enough provenance for credible reproduction
without violating data-use terms.
- After acceptance, replace anonymous archives with a public, licensed, citable artifact.
What AAMAS evidence reviewers open first
The single fact that shapes packaging: a reviewer will re-run a small **game** far sooner than they will retrain a large policy, so make the strategic core turnkey before polishing anything.
| Claim type | First artifact inspected | Common failure caught | |---|---|---| | Convergence to an equilibrium | The game definition and the learning-rule code | Solution concept named in the paper but not encoded in the evaluation | | Emergent cooperation/defection | The environment and reward specification | Result depends on an undocumented reward-shaping constant | | Beats other agents | The opponent/population set and match protocol | Only self-play reported; no held-out opponents | | Mechanism is truthful | The payment rule plus a strategic-deviation test | No script that lets an agent try to game the mechanism |
Worked vignette: packaging a self-play study
A hypothetical submission claims a learning rule that converges to a correlated equilibrium in a repeated congestion game, shown by self-play.
- Ship the game as one parameterized generator (number of agents, capacity, payoff scale)
rather than constants buried in a notebook, so reviewers can vary the interaction.
- Record the exact seed sequence and replication count behind every convergence plot; an
equilibrium-convergence claim without seeds is unfalsifiable.
- Emit payoff and regret tables directly from logged results so PDF and artifact numbers cannot
drift.
- Include a strategic-deviation harness: a script that drops in a non-conforming agent and
measures whether it profits, because that is exactly what a game-theory reviewer will try.
Calibration anchors
- Supplement inspection at AAMAS is at reviewer discretion; assume only the README and one entry
script get opened, and design the top level accordingly.
- Supplement size and format caps vary by cycle (25 MB single zip in 2026); verify against the
current OpenReview form rather than a past year.
Output format
[Artifact role] anonymous supplement / camera-ready release / public archive [Contents] <game/env/opponents/seeds/proofs/logs> [Anonymity risks] <paths/metadata/licenses/URLs> [Reproduction level] turnkey / scripted / descriptive / weak [Fixes before upload] <ordered list>
Stanford REAP × CoPaper.AI · 由斯坦福实证方法论团队精选与维护 访问 copaper.ai 微信:CoPaper.AI 按 11 个主流学科板块覆盖 经管与商科 社会科学 人文学科 数学与物理科学 生命科学 医学与健康 工程与技术 计算机科学与 AI 体育科学 点击任一学科名可跳转到对应说明;每类下的代表子领域在正文总览中完整列出。下方封面墙按 venue 导航,完整分类见覆盖一览。 🧭 布局指南 · 📚 Skill Pack 一览 · ⚡ 如何使用 · 🧪 自动实证
Other skills on awesome-journal-skills.
- /aaai-artifact-evaluation
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without violating double-blind or immutable-supplement rules.
Open skill - /aaai-author-response
Use when drafting an AAAI author response (rebuttal) under the single short character-limited author-feedback window, the no-URL rule, no-new-results guidance, AI-generated-review handling, and the AAAI two-phase review process where Phase-2 papers receive one feedback round
Open skill - /aaai-camera-ready
Use when preparing an accepted AAAI paper for camera-ready source submission to AAAI Press, including proceedings page limits, two-column template compliance, copyright transfer, purchased extra technical pages, deanonymization, registration, oral or poster presentation, and
Open skill - /aaai-experiments
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human evaluation, AI-for-Social-Impact and alignment/safety evidence, compute and cost reporting, and
Open skill - /aaai-related-work
Use when positioning an AAAI paper's novelty against archival work, contemporaneous arXiv or workshop papers, and AAAI/IJCAI/NeurIPS/ICML/ICLR neighbors across the broad AI scope, while staying inside AAAI's dual-submission and AI-as-source policy constraints and writing a
Open skill - /aaai-reproducibility
Use when strengthening an AAAI paper's reproducibility checklist (placed after references), experimental traceability, seed and hyperparameter reporting, compute and cost disclosure, dataset access and licensing, code/data ZIP readiness, and the claim-to-evidence map that
Open skill

