Skip to content
Content
Skill

/higgsfield-audio

Use when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance 1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok

From plugin
higgsfield-ai-prompt-skill
27632 skills2 commands
Install
$ npx -y skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-audio --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/higgsfield-audio

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance 1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok

SKILL.md

higgsfield-audio.SKILL.md
name: higgsfield-audio
description: >
  Use when the user asks about audio in Higgsfield videos, needs to add dialogue
  or lip-sync, wants sound effects or ambient sound in generated video, asks about
  music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance
  1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok Imagine Video). Also use when the user's
  prompt would benefit from audio direction but they haven't mentioned it.
  Also use when the user wants standalone audio — a soundtrack, ambience bed,
  multi-speaker scene audio (Seed Audio 1.0), or text-to-speech voiceover.
user-invocable: true
metadata:
  tags: [higgsfield, audio, dialogue, lip-sync, SFX, ambient, sound, BGM, music, voice, seed-audio, scene-audio, TTS]
  version: 3.5.0
  updated: 2026-08-09
  parent: higgsfield

Higgsfield Audio Prompting Guide

QUICK FACTS

*Routing aids — read the linked sections for the full rules.*

  • Native-joint audio models: Kling 3.0, Seedance 2.0 / 1.5 Pro, Veo 3/3.1, Grok — all others add audio in post [→](#which-models-support-audio)
  • Four layers to consider per prompt: Dialogue / SFX / Ambient / BGM [→](#the-four-audio-layers)
  • Lip-sync is the most failure-prone feature: 3–8s clips, MCU framing, one speaking face, locked camera, no head-motion tokens; per-language sync-word budgets are FIELD-reported [→](#lip-sync-rules)
  • **Seedance 2.0 `@Audio1` is a conditioning INPUT** — beat sync, the `[AUDIO: Xs]` script block, and the first-15s extraction trap [→](#audio-as-a-conditioning-input-seedance-20-audio1)
  • Multi-clip assembly: one master track · cuts land on musical punctuation, never inside a sung vowel (ECU mouth-match is the one exception) · unified grain + LUT masks batch color drift [→](#cutting-to-music-assembling-separately-generated-clips-on-one-track)
  • Cinema Studio 3.0 native joint audio (SCELA): describe audio as a separate section; specific foley beats generic moods [→](#cinema-studio-30-audio-businessteam-plan)
  • **Seed Audio 1.0** (`seed_audio`, standalone) = whole-scene audio in ONE pass — multi-speaker dialogue + music + SFX + ambience mixed [→](#scene-audio-generation-seed-audio-10)
  • Standalone Audio catalog (2026-08-01 snapshot): `seed_audio`, `qwen_audio_tts` (NEW — Qwen 3.0 TTS Flash, expressive instructions + cloned voices), `text2speech_v2` (5 engines incl. cozy_voice), plus 3 game-pipeline-only tools — distinct from in-video joint audio [→](#standalone-audio-tab-tool-catalog-2026-08-01-snapshot)

Which Models Support Audio?

| Model | Audio type | Dialogue | SFX | Ambient | BGM | Lip-sync | |-------|-----------|----------|-----|---------|-----|----------| | Kling 3.0 / Omni | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language | | Seedance 2.0 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language | | Seedance 1.5 Pro | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Best lip-sync | | Veo 3 / 3.1 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ English best | | Grok Imagine Video | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ | | All other models | ❌ | — | — | — | — | — |

**"Native joint"** means audio and video are generated simultaneously in one pass — not layered on after. This produces natural synchronization without post-production.

Models without native audio: add audio in post with Lipsync Studio or external tools.

---

The Four Audio Layers

Every audio-capable prompt should consider four layers. You don't need all four in every prompt, but knowing which to include gives the model clear direction.

1. Dialogue — What characters say

Put dialogue in quotes. Be explicit about who speaks, their tone, and language.

She says: "We need to leave. Now."
He whispers: "Not yet."

**Best practices:**

  • Keep dialogue short — 1-2 sentences per character per shot
  • Specify emotional tone: "says urgently", "whispers", "shouts across the room"
  • For non-English: specify language and dialect → `She speaks in Cantonese: "走啦"`
  • For Seedance 1.5 Pro: supports English, Chinese (incl. Sichuanese, Cantonese,

Taiwanese Mandarin, Shanghainese), Japanese, Korean, Spanish, Indonesian

2. SFX — Specific sound events tied to action

Describe SFX at the point they happen. Tie them to visible actions.

The glass shatters on the floor — sharp crack, then settling tinkle.
Footsteps on wet concrete — splashing, rhythmic.
A door slams shut — heavy metal, echoing.

**Best practices:**

  • One SFX description per action beat
  • Use onomatopoeia sparingly — descriptive phrases work better than "BANG" or "CRASH"
  • Tie timing to action: "as she sets the cup down" not "cup sound at 4 seconds"

3. Ambient — Background soundscape

Set the acoustic environment. This is the continuous sound bed.

Ambient: quiet café murmur, espresso machine, rain against windows.
Ambient: forest at night — crickets, distant owl, gentle wind through leaves.
Ambient: busy intersection — traffic, horns, construction in the distance.

**Best practices:**

  • 2-3 ambient elements maximum — more gets muddy
  • Describe the *space* acoustics: "reverberant church hall", "tight car interior"
  • Contrast silence with sound for impact: "Dead silence. Then — a single footstep."

4. BGM — Background music mood

Don't name songs or artists (content filter). Describe the musical texture.

BGM: slow piano, minor key, melancholic.
BGM: tense orchestral build — low strings, rising.
BGM: lo-fi hip-hop beat, warm vinyl crackle, relaxed.

**Best practices:**

  • Describe instrumentation, tempo, mood — not genre labels alone
  • "Tense strings, building" works better than "suspenseful music"
  • Specify when music enters/exits: "Piano enters at the midpoint, builds to the end"
  • For beat-sync content: "Cuts match the downbeat" or "Movement peaks on the drop"

---

Audio Prompt Structure

Add audio cues naturally within your prompt or as a dedicated block at the end.

Inline method (preferred for short prompts):

A woman walks into a quiet library. Her heels click on the marble floor — each step
echoing. She
Read more
Ships withhiggsfield-ai-prompt-skill

A comprehensive Claude skill library for generating high-quality prompts on Higgsfield AI — the cinematic video and image generation platform.

Get the whole plugin