Skip to content
Content
Skill

/video-gen

AI video generation via Seedance 2.0, Kling, and MiniMax Hailuo. Use when the user wants to generate a video clip — text-to-video, image-to-video, first/last-frame transitions, reference-guided generation, multi-shot, or generatively editing / extending an existing clip.

From plugin
openchatcut
91227 skills1 MCP
Install
$ npx -y skills add 0xsline/OpenChatCut --skill video-gen --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/video-gen

Context preview

The summary Claude sees to decide when to auto-load this skill.

AI video generation via Seedance 2.0, Kling, and MiniMax Hailuo. Use when the user wants to generate a video clip — text-to-video, image-to-video, first/last-frame transitions, reference-guided generation, multi-shot, or generatively editing / extending an existing clip.

SKILL.md

video-gen.SKILL.md
name: video-gen
description: |
  AI video generation via Seedance 2.0, Kling, and MiniMax Hailuo. Use when the user wants to generate a video clip — text-to-video, image-to-video, first/last-frame transitions, reference-guided generation, multi-shot, or generatively editing / extending an existing clip.
user-invocable: true

Video Gen

Submits one video generation job per call and returns a `jobId`. Job management (wait / status) belongs to `track_progress`; this skill does **not** place videos on the timeline automatically.

When to Use

Any time the user wants to generate a video clip — text-to-video, image-to-video, first-last-frame transition, reference-based generation, multi-shot storyboard, or generatively editing / extending an existing video (producing new generated footage based on a source clip; not timeline trimming).

Models

| Model | Reference | Strengths | | --- | --- | --- | | `seedance2` | [references/seedance2.md](references/seedance2.md) | Default when configured. Multimodal refs, first/last, edit/extend/bridge, 2–15s, 480p/720p/1080p/4k, audio/seed/camera/watermark/last-frame/task controls. | | `kling` | [references/kling.md](references/kling.md) | Technical camera/performance; Omni multi-shot; images ≤7 (≤4 with one feature `refVideos`); std/pro; 3–15s. | | `hailuo` | [references/hailuo.md](references/hailuo.md) | MiniMax 海螺. T2V / I2V / first+last; **6s or 10s**; 512P (Hailuo-02), 720p→768P, 1080P (6s); no multi-ref / multi-shot. |

**IMPORTANT:** Before generating, READ the chosen model's reference for capabilities, input channels, modes, prompt structure, and model-specific behavior. Never invent params the reference forbids.

Model Selection

Respect **configured vendors** from the capabilities prompt (only call a model whose key is on).

1. **User named a vendor** ("用海螺", "MiniMax", "Kling", "Seedance") → that `model`, if configured. 2. Else **default `seedance2`** when Seedance is configured. 3. Else if only Kling is on → `kling`. Else if only MiniMax is on → `hailuo`. 4. Switch away from default when:

  • Need **multi-shot customize / intelligence** → `kling` (confirm if not user-named).
  • Need **rich multi-modal refs** (video/audio refs, edit/extend) → `seedance2`.
  • Need a **short single beat** and only MiniMax is available, or user wants Hailuo → `hailuo` with duration 6 or 10.

If the required model is **not configured**, say so and offer: another configured video vendor, upload, or Motion Graphic — do not pretend the API exists.

Briefly tell the user what you will generate before submitting.

Tool Params

| Param | Values | Default | | --- | --- | --- | | `prompt` | video description | required (except Kling customize → use `multiPrompts`) | | `model` | `seedance2`, `kling`, `hailuo` | seedance2 when available | | `durationSeconds` | model-specific | seedance/kling ~5; **hailuo 6 or 10** (1080p → 6 only) | | `ratio` | see model docs | 16:9 (seedance/kling); **ignored on hailuo** | | `resolution` | `480p`, `512p`, `720p`, `1080p`, `4k` | provider-specific; hailuo adds 512p for Hailuo-02 | | `refVideoMode` | `feature`, `base` | kling only, with `refVideos` | | `promptOptimizer` / `fastPretreatment` | boolean | hailuo only | | `generateAudio`, `seed`, `cameraFixed`, `watermark` | controls | seedance only | | `returnLastFrame`, `executionExpiresAfter`, `priority` | controls | seedance only; requested last frame becomes another image asset | | `name` | descriptive asset name | required for good pool UX | | `firstFrame` | project image asset ref | optional | | `lastFrame` | project image asset ref | seedance / kling / hailuo (requires firstFrame; not with multi-ref on seedance) | | `refImages` / `refVideos` / `refAudios` | asset refs | seedance full; kling: images + **1** feature video (no audio); hailuo: none (frames / S2V subject) | | `mode` / `shotType` / `multiPrompts` | Kling multi-shot | kling only |

Model-specific params — see the model's reference.

Input Resolution

`firstFrame` / `lastFrame` / `refImages` / `refVideos` / `refAudios` all take a project asset reference. Prefer a full UUID or short prefix from `read_project`; `asset://<id>` and same-project asset URLs returned by `read_project` are also accepted. Per-slot type: frame slots and `refImages` → image; `refVideos` → video; `refAudios` → audio.

External URLs and base64 are not accepted. If the source is a public URL, download it into the project first (`download_media` for video/audio, `submit_image` for images) and pass the resulting asset id.

Workflow

Four-step loop. For each new generation, restart from Step 1 if the user's intent has shifted.

Step 1 — Align scope with the user

Before writing any prompt, align on three dimensions:

1. **Duration & segments** — total length, how many shots, and whether they live in one clip or several.

If the user has already stated a direction ("做一段", "in one video", "分别生成", "split into N shots", etc.), follow it — don't second-guess.

Otherwise, surface the two paths and let the user pick:

  • **Multi-shot within one clip** (see model ref) — single inference, subject / lighting / style physically consistent across sub-shots; fits a coherent narrative within the per-clip duration cap.
  • **Multiple clips** — each clip is independently controllable and re-rollable, but identity and style continuity have to be carried by anchors; fits durations beyond the cap or hard scene breaks.

Offer the trade-off; do not pick for the user.

2. **Content** — what each clip depicts. Summarize back what you understood, segment by segment. When content is vague (e.g. "generate a video of a girl dancing"), the user typically hasn't specified one or more of:

  • **Subject**: who / what is the main subject (appearance, outfit, defining features)?
  • **Action**: what are they doing? (For talking / emotional shots, what micro-expression?)
  • **Scene**: where — setting, time of day, environmental details?
  • **Ligh
Read more
Ships withopenchatcut

Open-source, local-first conversational AI video editor with a professional multi-track timeline, Agent Skills, MCP integration, and Remotion rendering.

Get the whole plugin

Other skills on openchatcut.