video-shotcraft
AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template
Topic in, narrated explainer video out. anything2explainer is a Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar, in Chinese or English.
$ npx -y skills add vincentwei1021/anything2explainer --agent claude-code
Repo: vincentwei1021/anything2explainer
What's inside
English | 简体中文
Topic in, narrated explainer video out. anything2explainer is a Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar, in Chinese or English. Every frame is drawn in code with Remotion (React + TypeScript). No stock footage, no generative video model, no frames lifted from anyone else's work.
It is not a CLI. What ships here is the whole method an AI coding agent needs to finish the film: a compilable Remotion template, a primitives and lighting library, tooling for voiceover / storyboard / rendering / quantitative QC, written style and motion specs, a multi-agent division-of-labour protocol, and one complete reference film as the quality bar.
English cut — RAG & Knowledge Bases, 5′02″, 44 lines / 785 words, voiced by kokoro-82m am_liam at natural speed:
https://github.com/user-attachments/assets/e2771c68-a28c-4459-ac5a-a5b685181eeb
Chinese cut — RAG 与知识库 v2, 4′54″, 44 lines / 1490 characters, dot-field backdrop (bg: 'dots'), voiced through the bring-your-own-TTS path (Volcengine TTS 2.0 + forced alignment):
https://github.com/user-attachments/assets/5c213990-cbba-439e-8371-fbb3aa348e05
Both cuts share one storyboard and 44 shots; the English cut re-times every shot to the English voiceover. The full paper trail of the original Chinese cut (4′35″, star-field backdrop, 8 build agents in parallel for 40 minutes, two QC rounds) lives in examples/rag/ (research → narration → storyboard → shot source → QC reports → delivery notes); rendered frames are in examples/rag/frames/.
| Frame / rate | 1280×720 @ 30fps, H.264 |
| Length | your call (see table below); 2–8 minutes all work |
| Language | Chinese or English (lang in src/config.ts); typography, subtitle budgets and TTS switch with it |
| Look | black canvas with one of two backdrops, star field + fog gradient or dot-field wave (bg in src/config.ts; the dot-field wave is ported from video-talkcraft); white line art + purple accents; ultra-bold headline type |
| Persistent layers | 44px white-on-black-stroke subtitles, bottom chapter progress bar, top capsule HUD, optional pipeline rail, built by Anything2Explainer skill end credit (builtBy, set to '' to drop) |
| Voiceover | Chinese: edge-tts zh-CN-YunxiNeural (Yunxi, male, unmodified rate ≈5.5 chars/s). English: kokoro-82m am_liam (Liam, male). Or bring your own TTS / finished audio |
Length drives how much ground the film covers, and the size of the whole pipeline:
| Length | Chinese chars | English words | Lines / shots | Build agents | Wall clock | Disk |
|---|---|---|---|---|---|---|
| 2–3 min | 650–880 | 280–420 | 24–32 | 4–6 | ≈1 h | ≈2 GB |
| 3–5 min (reference tier) | 1100–1400 | 420–700 | 40–50 | 8 | ≈2 h | ≈2 GB |
| 5–8 min | 1650–2200 | 700–1150 | 60–80 | 10–14 | ≈2–3 h | ≈3 GB |
Chapter count follows the content, within limits set by length: under 3 minutes use a single chapter (no chapter cards), 3–5 minutes 3–4 chapters of at least 60 s each, 5–8 minutes 4–6. The progress bar splits evenly across however many chapters the narration declares. A blank line in the narration marks a paragraph, which is also one shot: sentences inside a paragraph are separated by 10 frames, paragraph ends by 30, so the pause lands where the picture changes and every shot holds 1–1.5 s after its last element lands. The finished video runs 5–8% longer than the raw speech by design.
git clone https://github.com/Vincentwei1021/anything2explainer.git
ln -s "$PWD/anything2explainer" ~/.claude/skills/anything2explainer # Claude Code
ln -s "$PWD/anything2explainer" ~/.codex/skills/anything2explainer # Codex
Dependencies:
# Node ≥18 (the template's npm install pulls remotion 4.0.507 / react 19)
brew install ffmpeg # frame extraction / transcoding, required
python3 -m venv ~/.venvs/a2e && source ~/.venvs/a2e/bin/activate
pip install 'edge-tts==7.2.8' numpy pillow scipy # pin edge-tts: it tracks a Microsoft endpoint and breaks across upgrades (7.2.0+ needs word boundaries requested explicitly; the script does)
# only needed for English narration (kokoro-82m runs locally)
pip install kokoro soundfile && brew install espeak-ng
scipy is only used by the QC script frame_metrics.py. The shell scripts are zsh + Python 3, developed and verified on macOS; Linux should work, Windows is untested.
Verified on a Raspberry Pi 5 (ARM64, Python 3.13). Three things differ from macOS:
sudo apt install zsh espeak-ng # scripts are #!/bin/zsh; espeak-ng for kokoro/piper G2P
# Remotion has no linux-arm64 headless browser → point it at system Chromium:
sudo apt install chromium # or chromium-browser
export REMOTION_BROWSER_EXECUTABLE=/usr/bin/chromium # read by template/remotion.config.ts (no-op on macOS)
TTS on Linux/ARM. kokoro (the default English engine) is hard to install on ARM/Python 3.13 (it pins an old numpy and pulls spaCy → blis, which lack aarch64 wheels). Two local engines that install cleanly instead — pass one via TTS_ENGINE:
# kokoro_onnx — natural voice, onnxruntime (no torch/spaCy). Download model + voices from
# github.com/thewh1teagle/kokoro-onnx releases (kokoro-v1.0.onnx, voices-v1.0.bin)
pip install kokoro-onnx
TTS_ENGINE=kokoro_onnx KOKORO_ONNX_MODEL=…/kokoro-v1.0.onnx KOKORO_ONNX_VOICES=…/voices-v1.0.bin \
KOKORO_ONNX_VOICE=am_michael python3 scripts/tts_build.py
# piper — fastest local, robotic; a Pi-native fallback. Voice .onnx from github.com/rhasspy/piper
pip install piper-tts
TTS_ENGINE=piper PIPER_MODEL=…/en_US-ryan-medium.onnx python3 scripts/tts_build.py
The edge engine (natural, free, word-boundary timing) also works on Linux and needs no local model — it's a cloud call to Microsoft: TTS_ENGINE=edge VOICE=en-US-AndrewNeural python3 scripts/tts_build.py.
In Claude Code or Codex, just say what you want. The skill triggers itself:
Make me an explainer video about vector databases.
讲一下向量数据库,做成一条讲解视频
It then walks the 9 stages in SKILL.md:
You can also drive the template by hand:
template/scripts/new_project.sh ~/work/my-video myslug
cd ~/work/my-video
# 1. research/调研.md 2. script/narration.txt → python3 scripts/tts_build.py
# 3. script/storyboard_src.md → python3 scripts/render_storyboard.py 4. edit src/config.ts
# 5. src/shots/G1..Gn 6. scripts/preview.sh 30 (first 30 seconds)
# 7. VER=v1 scripts/render.sh + python3 scripts/frame_metrics.py 8. QC → fix → v2/v3
The run stops and waits for you at exactly four points instead of ploughing through (details in SKILL.md):
lang in src/config.ts, which drives typography, subtitle budgets and the default voice.| Tool class | What it produces | Where anything2explainer differs |
|---|---|---|
| Generative video models (Sora, Veo, Runway) | Footage synthesized from a prompt | Deterministic code, not pixels. Every number on screen traces to a source URL, and any frame can be fixed by editing one shot file |
| Avatar / presenter tools (HeyGen, Synthesia) | A digital presenter reading a script | No presenter. Motion-graphics diagrams that show the mechanism, with the narration driving the visuals |
| Remotion or Motion Canvas by hand | A programmable video canvas | Ships the method on top of the canvas: research → narration → storyboard → parallel build → QC, with style specs, motion vocabulary and a reference film to match |
| Manim | Python mathematical animations | An agent-driven end-to-end pipeline with TTS-aligned subtitles, chapters and QC; React / TypeScript rather than Python |
Which AI coding agents does it work with?
It is written for Claude Code and Codex, and those two are what it has been run with. The skill itself is plain Markdown plus a Remotion project, so any agent that reads SKILL.md-style skill folders and can run shell commands should be able to follow it.
Does it need a GPU? No. Remotion renders through headless Chromium on the CPU. The Chinese default voice (edge-tts) is a cloud call to a Microsoft endpoint; the English default (kokoro-82m) is an 82M-parameter model that runs locally on CPU.
Can I use my own voice or a different TTS?
Yes. Put the finished audio at public/assets/<slug>/audio.wav and fill src/common/timeline.ts and subs.ts by hand (format documented at the top of tts_build.py). Everything downstream is unchanged.
Can I change the visual style?
AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template
FAQ
anything2explainer is a Claude Code plugin with 1 hand-picked skill for video & audio work, indexed on Flowy. Install it with the command on its page. It includes anything2explainer. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it