ai-image-gen
通用 AI 生图:文生图 / 图生图 / 图像变体。当用户说 AI 生图、AI 画图、文生图、图生图、生成图片、生成配图、图像生成、AI 出图、AI…
从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录。 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。
$ npx -y skills add zju-real/easel --skill clipify --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/clipifyContext preview
The summary Claude sees to decide when to auto-load this skill.
从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录。 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。
name: clipify description: >- 从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录。 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。 layer: produce
Find the funniest moments in a video, cut them as standalone clips, optionally reformat 16:9 → 9:16 (face-pan or split-screen), and burn opus-style word-by-word captions.
Working dir: `/tmp/clipify/` (mkdir at start, leave artifacts for debugging).
---
mkdir -p /tmp/clipify ffmpeg -y -i "$VIDEO" -vn -ac 1 -ar 16000 /tmp/clipify/audio.wav whisper /tmp/clipify/audio.wav --model tiny.en --word_timestamps True --output_format json --output_dir /tmp/clipify --language en
Read the resulting JSON (or `.txt`) and pick 3–5 candidate clips. Funny signals to scan for:
For each candidate, propose: `[start, end, why-it's-funny, suggested title]`. Aim for 10–25s clips. Show the list and let the user confirm/pick.
ffmpeg -y -ss "$START" -t "$DURATION" -i "$VIDEO" -c copy /tmp/clipify/clip_$N.mp4
(Use `-c copy` for instant trim. Re-encode only if cuts must be frame-accurate.)
Ask the user (skip if they already specified): "9:16 (TikTok / Reels), 16:9 (YouTube), or 1:1 (Insta feed)?"
Detect source aspect with `ffprobe`. If source is 16:9 and target is 9:16, ask:
> "Two options: **(a) hard-cut pan** that follows whoever is speaking (single face on screen at a time), or **(b) split-screen** stack with both faces visible. Which do you want?"
Skip the question if there's only one face (single-talker clip). For single-talker, just center-crop.
1. **Locate the two face ROIs.** Sample one frame: `ffmpeg -ss <middle> -i <clip> -frames:v 1 /tmp/clipify/probe.jpg`. Read it. Eyeball each face's mouth+chin area as `x,y,w,h` in the source's pixel space. (No cv2 needed — camera is static within a clip; one frame is enough.) Verify by drawing boxes:
ffmpeg -i probe.jpg -vf "drawbox=x=$LX:y=$LY:w=$LW:h=$LH:color=cyan@0.9:t=4,drawbox=x=$RX:y=$RY:w=$RW:h=$RH:color=magenta@0.9:t=4" verify.jpg
Iterate **at most twice**. Boxes should cover mouth + chin and avoid hands/mics. Don't over-tune — frame differencing is forgiving.
2. **Extract per-frame motion energy in each ROI:**
ffmpeg -y -i clip.mp4 -filter_complex " [0:v]split=2[a][b]; [a]crop=$LW:$LH:$LX:$LY,format=gray,tblend=all_mode=difference,signalstats,metadata=mode=print:key=lavfi.signalstats.YAVG:file=/tmp/clipify/L.txt[la]; [b]crop=$RW:$RH:$RX:$RY,format=gray,tblend=all_mode=difference,signalstats,metadata=mode=print:key=lavfi.signalstats.YAVG:file=/tmp/clipify/R.txt[ra] " -map "[la]" -f null - -map "[ra]" -f null -
3. **Build speaker timeline** (min dwell 1.0s — short interjections merge into the prior speaker):
python3 <skill-dir>/scripts/analyze.py /tmp/clipify/L.txt /tmp/clipify/R.txt 1.0 > /tmp/clipify/segments.json
4. **Pick pan x-coordinates** for a 9:16 vertical strip from the source. With source W=1920 and target W=1080, crop strip width = 608.
5. **Generate the hard-cut x expression and render:**
EXPR=$(python3 <skill-dir>/scripts/build_pan.py /tmp/clipify/segments.json $LEFT_X $RIGHT_X)
ffmpeg -y -i clip.mp4 -filter_complex \
"[0:v]crop=608:1080:x='$EXPR':y=0,scale=1080:1920:flags=lanczos[v]" \
-map "[v]" -map 0:a -c:v libx264 -preset fast -crf 20 -pix_fmt yuv420p \
-c:a aac -b:a 192k /tmp/clipify/clip_panned.mp4Source 1920×1080 assumed; for 4K source either downscale first or double all coordinates.
Two stacked tiles, 1080×960 each. The active speaker's tile is on top — overlay flips at speaker changes.
[0:v]split=2[a0][a1]; [a0]crop=Wcrop:Hcrop:LX_tile:LY_tile,scale=1080:960,split=2[lt0][lt1]; [a1]crop=Wcrop:Hcrop:RX_tile:RY_tile,scale=1080:960,split=2[rt0][rt1]; [lt0][rt0]vstack
An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的 AI 社交媒体智能体——发现热点趋势、创作内容、一键发布至各大平台,并学习分析哪些内容真正有效,覆盖小红书、抖音、知乎、哔哩哔哩等平台。
Repo: zju-real/easel
通用 AI 生图:文生图 / 图生图 / 图像变体。当用户说 AI 生图、AI 画图、文生图、图生图、生成图片、生成配图、图像生成、AI 出图、AI…
AI 音乐 / BGM 生成:给短视频、社媒内容生成原创背景音乐 / 配乐 / 纯音乐。通过可插拔 provider(阿里 DashScope / Suno 类第三方…
AI 视频生成:文生视频 / 图生视频 / 数字人首帧驱动。通过可插拔 provider(通义万相 Wan / 火山 Seedance / 快手可灵 / OpenAI…
outputs/ 目录下的产物管理:按日期/平台/类型归档、打标签、搜索历史内容、生成素材清单。 当用户说"整理素材"、"归档"、"找之前的内容"、"搜索历史"、"素材管理"、…
音频降噪:去除录音中的背景噪声、电流声、风噪、嗡嗡声,基于 ffmpeg 滤镜链(afftdn/highpass/lowpass)。…
通用音频处理:音频剪辑/裁剪、格式转码(mp3/wav/m4a/aac)、音量归一化、从视频提取音轨、多段拼接、淡入淡出、变速(保音高)。当用户说“剪音频”“裁一段”“转成…