/acmmm-writing-style
Use when revising an ACM MM (ACM Multimedia) paper for house style — putting the cross-modal contribution on the first page, framing media (figures, video, audio) as evidence rather than decoration, making the fusion the visible claim, and compressing the argument into a 6-8
$ npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill acmmm-writing-style --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/acmmm-writing-style
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when revising an ACM MM (ACM Multimedia) paper for house style — putting the cross-modal contribution on the first page, framing media (figures, video, audio) as evidence rather than decoration, making the fusion the visible claim, and compressing the argument into a 6-8
SKILL.md
acmmm-writing-style.SKILL.mdname: acmmm-writing-style
description: Use when revising an ACM MM (ACM Multimedia) paper for house style — putting the cross-modal contribution on the first page, framing media (figures, video, audio) as evidence rather than decoration, making the fusion the visible claim, and compressing the argument into a 6-8 page ACM sigconf body with references-only overflow.
ACM MM Writing Style
Use this to revise an ACM Multimedia draft for the venue's expectations. The reader is a busy reviewer scanning across sixteen thematic areas; the paper must announce *what is multimedia about it* before the model diagram.
The ACM MM first-page arc
Lead the abstract and first paragraph with: **problem → why one modality is insufficient → the cross-modal method or system → media-grounded evidence → why it matters for multimedia**. The multimedia contribution belongs on page one; a paper that opens with a single-modality benchmark reads as a CVPR or ACL paper that wandered in.
- State the **fusion or systems mechanism** as the contribution, not a backbone swap.
- Make every media element *do work*: a teaser figure that shows the cross-modal signal, a
video that is evidence for a claim, an audio clip that a reader can check.
- Scope claims to what the evidence supports; "improves engagement" needs a measured
engagement result behind it.
Media-as-evidence table
| Media element | Weak (decoration) | Strong (evidence) | |---|---|---| | Teaser figure | A pretty system diagram | The moment where modalities disagree and the method wins | | Qualitative grid | Cherry-picked successes | Paired success/failure across modalities with captions that state the point | | Supplementary video | "See our results" | A clip tied to a specific claim, with the baseline shown alongside | | Audio sample | An unlabeled waveform | The case the vision-only baseline misses, annotated |
Compression into the sigconf body
The 6–8 page body is short by ACM standards, and figures compete with text for space.
- Push proofs, full protocols, extra qualitative media, and hyperparameter tables to the
**supplement**; keep the body's argument self-contained without them.
- Write **self-contained captions** — a reviewer reading only figures and captions should
grasp the cross-modal claim.
- Budget page space before writing: decide which two or three media elements earn body space
and which move to the supplement.
Revision passes
Pass 1 (contribution): Does page one name the cross-modal/systems contribution?
Pass 2 (fusion): Is the mechanism the claim, and is it ablated later?
Pass 3 (media): Does each figure/clip support a specific claim, with a self-contained caption?
Pass 4 (scope): Is every "better/more engaging/higher quality" tied to a measured result?
Pass 5 (budget): Does the body fit 6-8 sigconf pages with overflow holding references only?
Title and abstract for cross-area reviewers
Your reviewers may come from different thematic areas, so the title and abstract have to be legible to a vision person, an audio person, and a systems person at once.
- Put the **modalities and the mechanism** in the title where natural ("audio-visual," "text-and-
image," "cross-modal") so area chairs assign the right reviewers.
- Make the abstract's first two sentences carry the whole contribution; a reviewer triaging many
papers may read little more.
- Avoid single-community jargon in the abstract; define the one term your cross-area readers will
not share.
A body-budget worked pass
Treat the 6–8 page limit as a budget you allocate before writing prose:
p1 intro: contribution + why one modality fails + teaser figure
p2 related work (tight) + problem setup
p3-4 method: the fusion/alignment mechanism, one architecture figure
p5-6 experiments: main table, the decisive ablation, failure cases, user-study summary
p7-8 discussion + limitations; references spill onto the overflow pages (references only)
If a section will not fit, move detail to the supplement rather than shrinking the font or margins — template tampering is a desk-reject risk, and a cramped body reads worse than a clean one with a fuller supplement.
Common ACM MM style failures
- **Modality as garnish** — audio/text mentioned but never shown to matter; fix by leading
with the seam and ablating it.
- **Vision-paper voice** — the whole framing is a benchmark race; fix by foregrounding the
multimedia question.
- **Unbacked perceptual claims** — "more natural," "more engaging" with no user study; fix
by measuring or softening.
- **Caption starvation** — figures that only make sense from the body text; fix by making
captions stand alone.
- **Overflow abuse** — method text pushed onto the references pages; those pages are for
references only, and misuse risks desk reject.
Output format
[First-page verdict] multimedia contribution up front / buried
[Fusion visibility] mechanism is the claim / hidden behind a backbone
[Media evidence] each element earns its place / decorative elements: <list>
[Scope] claims matched to evidence / overclaims: <list>
[Page budget] fits 6-8 sigconf pages / over by <n>
[Top three fixes] <ordered>
Read more
name: acmmm-writing-style description: Use when revising an ACM MM (ACM Multimedia) paper for house style — putting the cross-modal contribution on the first page, framing media (figures, video, audio) as evidence rather than decoration, making the fusion the visible claim, and compressing the argument into a 6-8 page ACM sigconf body with references-only overflow.
ACM MM Writing Style
Use this to revise an ACM Multimedia draft for the venue's expectations. The reader is a busy reviewer scanning across sixteen thematic areas; the paper must announce *what is multimedia about it* before the model diagram.
The ACM MM first-page arc
Lead the abstract and first paragraph with: **problem → why one modality is insufficient → the cross-modal method or system → media-grounded evidence → why it matters for multimedia**. The multimedia contribution belongs on page one; a paper that opens with a single-modality benchmark reads as a CVPR or ACL paper that wandered in.
- State the **fusion or systems mechanism** as the contribution, not a backbone swap.
- Make every media element *do work*: a teaser figure that shows the cross-modal signal, a
video that is evidence for a claim, an audio clip that a reader can check.
- Scope claims to what the evidence supports; "improves engagement" needs a measured
engagement result behind it.
Media-as-evidence table
| Media element | Weak (decoration) | Strong (evidence) | |---|---|---| | Teaser figure | A pretty system diagram | The moment where modalities disagree and the method wins | | Qualitative grid | Cherry-picked successes | Paired success/failure across modalities with captions that state the point | | Supplementary video | "See our results" | A clip tied to a specific claim, with the baseline shown alongside | | Audio sample | An unlabeled waveform | The case the vision-only baseline misses, annotated |
Compression into the sigconf body
The 6–8 page body is short by ACM standards, and figures compete with text for space.
- Push proofs, full protocols, extra qualitative media, and hyperparameter tables to the
**supplement**; keep the body's argument self-contained without them.
- Write **self-contained captions** — a reviewer reading only figures and captions should
grasp the cross-modal claim.
- Budget page space before writing: decide which two or three media elements earn body space
and which move to the supplement.
Revision passes
Pass 1 (contribution): Does page one name the cross-modal/systems contribution? Pass 2 (fusion): Is the mechanism the claim, and is it ablated later? Pass 3 (media): Does each figure/clip support a specific claim, with a self-contained caption? Pass 4 (scope): Is every "better/more engaging/higher quality" tied to a measured result? Pass 5 (budget): Does the body fit 6-8 sigconf pages with overflow holding references only?
Title and abstract for cross-area reviewers
Your reviewers may come from different thematic areas, so the title and abstract have to be legible to a vision person, an audio person, and a systems person at once.
- Put the **modalities and the mechanism** in the title where natural ("audio-visual," "text-and-
image," "cross-modal") so area chairs assign the right reviewers.
- Make the abstract's first two sentences carry the whole contribution; a reviewer triaging many
papers may read little more.
- Avoid single-community jargon in the abstract; define the one term your cross-area readers will
not share.
A body-budget worked pass
Treat the 6–8 page limit as a budget you allocate before writing prose:
p1 intro: contribution + why one modality fails + teaser figure p2 related work (tight) + problem setup p3-4 method: the fusion/alignment mechanism, one architecture figure p5-6 experiments: main table, the decisive ablation, failure cases, user-study summary p7-8 discussion + limitations; references spill onto the overflow pages (references only)
If a section will not fit, move detail to the supplement rather than shrinking the font or margins — template tampering is a desk-reject risk, and a cramped body reads worse than a clean one with a fuller supplement.
Common ACM MM style failures
- **Modality as garnish** — audio/text mentioned but never shown to matter; fix by leading
with the seam and ablating it.
- **Vision-paper voice** — the whole framing is a benchmark race; fix by foregrounding the
multimedia question.
- **Unbacked perceptual claims** — "more natural," "more engaging" with no user study; fix
by measuring or softening.
- **Caption starvation** — figures that only make sense from the body text; fix by making
captions stand alone.
- **Overflow abuse** — method text pushed onto the references pages; those pages are for
references only, and misuse risks desk reject.
Output format
[First-page verdict] multimedia contribution up front / buried [Fusion visibility] mechanism is the claim / hidden behind a backbone [Media evidence] each element earns its place / decorative elements: <list> [Scope] claims matched to evidence / overclaims: <list> [Page budget] fits 6-8 sigconf pages / over by <n> [Top three fixes] <ordered>
Stanford REAP × CoPaper.AI · 由斯坦福实证方法论团队精选与维护 访问 copaper.ai 微信:CoPaper.AI 按 11 个主流学科板块覆盖 经管与商科 社会科学 人文学科 数学与物理科学 生命科学 医学与健康 工程与技术 计算机科学与 AI 体育科学 点击任一学科名可跳转到对应说明;每类下的代表子领域在正文总览中完整列出。下方封面墙按 venue 导航,完整分类见覆盖一览。 🧭 布局指南 · 📚 Skill Pack 一览 · ⚡ 如何使用 · 🧪 自动实证
Other skills on awesome-journal-skills.
- /aaai-artifact-evaluation
Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without violating double-blind or immutable-supplement rules.
Open skill - /aaai-author-response
Use when drafting an AAAI author response (rebuttal) under the single short character-limited author-feedback window, the no-URL rule, no-new-results guidance, AI-generated-review handling, and the AAAI two-phase review process where Phase-2 papers receive one feedback round
Open skill - /aaai-camera-ready
Use when preparing an accepted AAAI paper for camera-ready source submission to AAAI Press, including proceedings page limits, two-column template compliance, copyright transfer, purchased extra technical pages, deanonymization, registration, oral or poster presentation, and
Open skill - /aaai-experiments
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human evaluation, AI-for-Social-Impact and alignment/safety evidence, compute and cost reporting, and
Open skill - /aaai-related-work
Use when positioning an AAAI paper's novelty against archival work, contemporaneous arXiv or workshop papers, and AAAI/IJCAI/NeurIPS/ICML/ICLR neighbors across the broad AI scope, while staying inside AAAI's dual-submission and AI-as-source policy constraints and writing a
Open skill - /aaai-reproducibility
Use when strengthening an AAAI paper's reproducibility checklist (placed after references), experimental traceability, seed and hyperparameter reporting, compute and cost disclosure, dataset access and licensing, code/data ZIP readiness, and the claim-to-evidence map that
Open skill

