/voice-compose
Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generate_voice.js CLI. Use before calling generate_voice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for
$ npx -y skills add Utopai-Research/pai-pro --skill voice-compose --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/voice-compose
Context preview
The summary Claude sees to decide when to auto-load this skill.
Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generate_voice.js CLI. Use before calling generate_voice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for
SKILL.md
voice-compose.SKILL.mdname: voice-compose
description: Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generate_voice.js CLI. Use before calling generate_voice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for every speaking character or VO/narration, or create exact narration/VO/final line audio.
Default to one short reusable timbre sample per speaking character and one VO/narrator sample when narration exists. `video-compose` keeps actual shot dialogue/VO in the video prompt. Treat `audio_result.data.text` as downstream speech only for approved final narration/line reads.
Patterns
Follow `PROJECT_AGENT.md` for context/staging. This skill owns voice prompt + CLI shape.
1. Character voice sample
Triggers: "give / design a voice for [character]", "what does [character] sound like", "voices for all the characters on the canvas".
- Target: any `image_result` for the person; don't gate on subtype. Read `data.local_path` before prompt; layer `name`/`role`/`description` on top.
- Call:
node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \
--text "<line>" \
--prompt "<voice design brief>" \
--source-node-id <character.id>- Prompt describes the **voice**, not the character:
> `[age bracket] [gender], [timbre], [register], [pace], [accent if relevant]. [optional emotional color].`
✅ "Mid-50s man, gravelly baritone, measured pace, slight rasp from decades of smoking, weary but steady." ✅ "Young woman, bright mezzo, warm, quick and percussive. Slight Southern lilt." ❌ "Detective Morris's voice." — names the character, not the voice. The model needs sound qualities.
- `text`: 1-3 sentence in-character sample (≤200 chars), not every script line.
- Script breakdowns: one staged call per speaker; preserve labels. Add separate VO/narrator via Pattern 2.
2. Narrator / VO voice sample or final line audio
Triggers: narrator voice, voice-over, "a voice that says X" without character, narration track, or explicit final line-read audio.
- Omit `--source-node-id`:
node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \
--text "<the narration line>" \
--prompt "<voice design brief>"- Same prompt convention as Pattern 1.
- Reusable VO/narrator anchor: short sample line in narrator style, not full script narration.
- Final narration/line-read: copy approved text exactly into `--text`; then `data.text` is source of truth.
Read more
name: voice-compose description: Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generate_voice.js CLI. Use before calling generate_voice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for every speaking character or VO/narration, or create exact narration/VO/final line audio.
Default to one short reusable timbre sample per speaking character and one VO/narrator sample when narration exists. `video-compose` keeps actual shot dialogue/VO in the video prompt. Treat `audio_result.data.text` as downstream speech only for approved final narration/line reads.
Patterns
Follow `PROJECT_AGENT.md` for context/staging. This skill owns voice prompt + CLI shape.
1. Character voice sample
Triggers: "give / design a voice for [character]", "what does [character] sound like", "voices for all the characters on the canvas".
- Target: any `image_result` for the person; don't gate on subtype. Read `data.local_path` before prompt; layer `name`/`role`/`description` on top.
- Call:
node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \
--text "<line>" \
--prompt "<voice design brief>" \
--source-node-id <character.id>- Prompt describes the **voice**, not the character:
> `[age bracket] [gender], [timbre], [register], [pace], [accent if relevant]. [optional emotional color].`
✅ "Mid-50s man, gravelly baritone, measured pace, slight rasp from decades of smoking, weary but steady." ✅ "Young woman, bright mezzo, warm, quick and percussive. Slight Southern lilt." ❌ "Detective Morris's voice." — names the character, not the voice. The model needs sound qualities.
- `text`: 1-3 sentence in-character sample (≤200 chars), not every script line.
- Script breakdowns: one staged call per speaker; preserve labels. Add separate VO/narrator via Pattern 2.
2. Narrator / VO voice sample or final line audio
Triggers: narrator voice, voice-over, "a voice that says X" without character, narration track, or explicit final line-read audio.
- Omit `--source-node-id`:
node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \
--text "<the narration line>" \
--prompt "<voice design brief>"- Same prompt convention as Pattern 1.
- Reusable VO/narrator anchor: short sample line in narrator style, not full script narration.
- Final narration/line-read: copy approved text exactly into `--text`; then `data.text` is source of truth.
The local AI filmmaking studio, driven from your coding agent. ![Discord][discord-url] [][claude-code-url] [][codex-url]
Repo: Utopai-Research/pai-pro
Other skills on pai-pro.
- /groups-compose
Designs and maintains semantic groupings and readable layouts on the filmmaking canvas — scenes, character-reference sets, act beats, and other titled visual frames. Use when nodes on the canvas cluster around a shared meaning and would read more clearly if arranged together and
Open skill - /image-compose
Generates/edits filmmaking canvas images via generate_image.js and generate_image_pro.js. Use before image CLIs for character/location design, refs, starting frames, storyboards, stills, edits, variations, and downstream video anchors. Video-bound characters default to Pattern 7
Open skill - /script-compose
Handles explicit screenplay/story work on the filmmaking canvas. Triages screenplay (use verbatim), story/concept (iterate then rewrite), or neither (defer). Captures the final script note/title; on explicit command, splits into <=15s shot notes and extracts characters,
Open skill - /story-to-video-workflow
Orchestrates story, script, screenplay, concept, product promo, and multi-shot idea work into finished video. Use first when the user asks to make a video from a story or script; asks what next in a story video project; or needs a decision spanning script splitting, image refs,
Open skill - /video-compose
Generates and prompts video clips on the filmmaking canvas. Use when the user asks to generate, render, animate, continue, restyle, edit, shoot, or compose a video clip; render script or shot notes as video; animate a storyboard, starting frame, image, character, location, or
Open skill

