Skip to content
Content
Skill

/video-understand

Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without

From plugin
openmontage
46k48 skills5 agents3 commands
Install
$ npx -y skills add calesthio/OpenMontage --skill video-understand --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/video-understand

Context preview

The summary Claude sees to decide when to auto-load this skill.

Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without

SKILL.md

video-understand.SKILL.md
name: video-understand
description: |
  Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed.
  Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally,
  (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

video-understand

Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.

Prerequisites

  • `ffmpeg` + `ffprobe` (required): `brew install ffmpeg`
  • `openai-whisper` (optional, for transcription): `pip install openai-whisper`

Commands

# Scene detection + transcribe (default)
python3 skills/video-understand/scripts/understand_video.py video.mp4

# Keyframe extraction
python3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe

# Regular interval extraction
python3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval

# Limit frames extracted
python3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10

# Use a larger Whisper model
python3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small

# Frames only, skip transcription
python3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe

# Quiet mode (JSON only, no progress)
python3 skills/video-understand/scripts/understand_video.py video.mp4 -q

# Output to file
python3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json

CLI Options

| Flag | Description | |------|-------------| | `video` | Input video file (positional, required) | | `-m, --mode` | Extraction mode: `scene` (default), `keyframe`, `interval` | | `--max-frames` | Maximum frames to keep (default: 20) | | `--whisper-model` | Whisper model size: tiny, base, small, medium, large (default: base) | | `--no-transcribe` | Skip audio transcription, extract frames only | | `-o, --output` | Write result JSON to file instead of stdout | | `-q, --quiet` | Suppress progress messages, output only JSON |

Extraction Modes

| Mode | How it works | Best for | |------|-------------|----------| | `scene` | Detects scene changes via ffmpeg `select='gt(scene,0.3)'` | Most videos, varied content | | `keyframe` | Extracts I-frames (codec keyframes) | Encoded video with natural keyframe placement | | `interval` | Evenly spaced frames based on duration and max-frames | Fixed sampling, predictable output |

If `scene` mode detects no scene changes, it automatically falls back to `interval` mode.

Output

The script outputs JSON to stdout (or file with `-o`). See `references/output-format.md` for the full schema.

{
  "video": "video.mp4",
  "duration": 18.076,
  "resolution": {"width": 1224, "height": 1080},
  "mode": "scene",
  "frames": [
    {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"}
  ],
  "frame_count": 12,
  "transcript": [
    {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."}
  ],
  "text": "Full transcript...",
  "note": "Use the Read tool to view frame images for visual understanding."
}

Use the Read tool on frame image paths to visually inspect extracted frames.

References

  • `references/output-format.md` -- Full JSON output schema documentation
Read more
Ships withopenmontage

World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.

Get the whole plugin

Other skills on openmontage.