Skip to content
Development
Skill

/watch

Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by

BOOST
From plugin
armory
32887 skills1 agent1 command
Install
$ npx -y skills add Mathews-Tom/armory --skill watch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/watch

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by

SKILL.md

watch.SKILL.md
name: watch
description: 'Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by keyword (use youtube-search) or creating videos (use concept-to-video or remotion-video).'
metadata:
  version: 3.0.0
  category: visualization
  tags: [video, youtube, analysis, multimodal, evidence]
  difficulty: intermediate
  complements: [youtube-search, notebooklm, concept-to-video, remotion-video]

Watch

Analyze an existing video from timestamped speech and locally inspected visual evidence. Answer the user's question first; preserve structured concept analysis for general summaries. A transcript explains what was said, not everything shown. Watch processes media locally and does not upload video or audio to a media-analysis service.

When to Use

| Request | Evidence mode | Boundary | |---|---|---| | Summarize spoken ideas, an interview, or a podcast | `transcript` | No video download when captions suffice | | Inspect a slide, code, UI demo, or private recording | `local` | Read bounded frames and available speech evidence | | Search a long public video for a visual moment | `local` | Start with bounded sampling, then inspect focused intervals; coverage is not exhaustive | | Recreate a visual reference | `local` | Pass inspected evidence to a generation skill; Watch does not generate video | | Align supplied retention analytics with content | Focused `local` | Association is not proof of why viewers left |

Do not activate for video discovery, new-video creation, financial watchlists, or watching filesystem changes. Other URL sources use yt-dlp support, not a promise that every website or private video is accessible.

Prerequisites

`SKILL_DIR` is the absolute directory containing this file; scripts sit beside it. Resolve this path from the installed skill, not the current project directory. Python 3.12+ and uv are required. Local media inspection also needs FFmpeg/ffprobe. The documented uv invocation supplies caption/downloader dependencies without changing the user's project.

uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" --help
ffmpeg -version
ffprobe -version

Check only dependencies relevant to the selected path. Transcript-only requests do not require FFmpeg. Never install system binaries or large speech models without explicit user authorization. Local processing means the agent's execution machine, not automatically the user's laptop; captions and frames opened by the host assistant remain subject to that host's data policy.

Workflow

1. Select the question and evidence boundary

1. Preserve the user's question verbatim in `--question`. Without a question, produce a structured summary. 2. Select `--engine transcript` when speech alone answers the request. Select `--engine local` when visuals matter. Runtime default is local. 3. Keep media analysis local. Missing captions or speech are evidence gaps, not permission to upload media or select another service. 4. Choose quick, standard, or deep analysis using `--depth`. This controls presentation, not frame coverage. A deep summary is not an exhaustive visual inspection.

2. Acquire captions without unnecessary media

uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "YOUTUBE_URL" --engine transcript --depth standard --question "Explain the main ideas and actionable takeaways"

YouTube captions use youtube-transcript-api first, then a selected yt-dlp caption track. Other supported URLs use yt-dlp. Read the reported source, manual/automatic kind, actual language, and gaps. Do not claim a requested language was used when a different track was selected, or infer the speaker's language from translated captions.

The report retains source-relative segment timestamps. No captions is missing speech evidence, not evidence of silence. Local files require explicit speech transcription to obtain a transcript; a visual-only result remains useful.

3. Inspect bounded visual evidence

uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "URL_OR_LOCAL_FILE" --engine local --question "Identify the tool shown on screen" --detail balanced --max-frames 40 --resolution 1024

Read **every listed frame** using Read before claiming what is shown. Combine the images with the timestamped transcript. The report labels each frame with its actual decoded source time and selection reason. Fixed budgets produce sparse coverage on long recordings; do not turn a sampled absence into “never appears.”

  • `efficient`: keyframe selection with uniform fallback.
  • `balanced`: scene-aware selection with uniform fallback.
  • `transcript` detail under the local engine: captions plus explicitly requested cue frames only.
  • `--no-dedup`: preserve near-identical selected images when small text, code, cursor, or UI changes matter. Deduplication is not event detection.

Start with a bounded scan. Identify relevant speech cues (“look here,” “this diagram”) and visual candidates, then inspect a tighter interval or explicit timestamps. A visual event need not be mentioned in speech; do not use captions as the only search index.

uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "URL_OR_LOCAL_FILE" --engine local --start 02:00 --end 02:25 --detail transcript --timestamps 02:13 --max-frames 4 --no-dedup --question "Read the tool name and explain the demonstration"

Focus times and cues are absolute source times; seconds, MM:SS, and HH:MM:SS are accepted. Frames lie inside `[start, end)`. Requested cue time and actual decoded time are distinct. Cue frames reserve space in the total cap; an excessiv

Read more
Ships witharmory

Curated, production-grade skills, agents, hooks, rules, commands, utilities, and presets for AI coding agents. No magic, no demos — battle-tested workflows built for developers who use AI seriously.

Get the whole plugin

Other skills on armory.