adr-writer
Generates Architecture Decision Records capturing context, rationale, alternatives, and…
Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by
$ npx -y skills add Mathews-Tom/armory --skill watch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/watchContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by
name: watch description: 'Use when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what happens on screen", "extract concepts from video", or "video key points". NOT for finding videos by keyword (use youtube-search) or creating videos (use concept-to-video or remotion-video).' metadata: version: 3.0.0 category: visualization tags: [video, youtube, analysis, multimodal, evidence] difficulty: intermediate complements: [youtube-search, notebooklm, concept-to-video, remotion-video]
Analyze an existing video from timestamped speech and locally inspected visual evidence. Answer the user's question first; preserve structured concept analysis for general summaries. A transcript explains what was said, not everything shown. Watch processes media locally and does not upload video or audio to a media-analysis service.
| Request | Evidence mode | Boundary | |---|---|---| | Summarize spoken ideas, an interview, or a podcast | `transcript` | No video download when captions suffice | | Inspect a slide, code, UI demo, or private recording | `local` | Read bounded frames and available speech evidence | | Search a long public video for a visual moment | `local` | Start with bounded sampling, then inspect focused intervals; coverage is not exhaustive | | Recreate a visual reference | `local` | Pass inspected evidence to a generation skill; Watch does not generate video | | Align supplied retention analytics with content | Focused `local` | Association is not proof of why viewers left |
Do not activate for video discovery, new-video creation, financial watchlists, or watching filesystem changes. Other URL sources use yt-dlp support, not a promise that every website or private video is accessible.
`SKILL_DIR` is the absolute directory containing this file; scripts sit beside it. Resolve this path from the installed skill, not the current project directory. Python 3.12+ and uv are required. Local media inspection also needs FFmpeg/ffprobe. The documented uv invocation supplies caption/downloader dependencies without changing the user's project.
uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" --help ffmpeg -version ffprobe -version
Check only dependencies relevant to the selected path. Transcript-only requests do not require FFmpeg. Never install system binaries or large speech models without explicit user authorization. Local processing means the agent's execution machine, not automatically the user's laptop; captions and frames opened by the host assistant remain subject to that host's data policy.
1. Preserve the user's question verbatim in `--question`. Without a question, produce a structured summary. 2. Select `--engine transcript` when speech alone answers the request. Select `--engine local` when visuals matter. Runtime default is local. 3. Keep media analysis local. Missing captions or speech are evidence gaps, not permission to upload media or select another service. 4. Choose quick, standard, or deep analysis using `--depth`. This controls presentation, not frame coverage. A deep summary is not an exhaustive visual inspection.
uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "YOUTUBE_URL" --engine transcript --depth standard --question "Explain the main ideas and actionable takeaways"
YouTube captions use youtube-transcript-api first, then a selected yt-dlp caption track. Other supported URLs use yt-dlp. Read the reported source, manual/automatic kind, actual language, and gaps. Do not claim a requested language was used when a different track was selected, or infer the speaker's language from translated captions.
The report retains source-relative segment timestamps. No captions is missing speech evidence, not evidence of silence. Local files require explicit speech transcription to obtain a transcript; a visual-only result remains useful.
uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "URL_OR_LOCAL_FILE" --engine local --question "Identify the tool shown on screen" --detail balanced --max-frames 40 --resolution 1024
Read **every listed frame** using Read before claiming what is shown. Combine the images with the timestamped transcript. The report labels each frame with its actual decoded source time and selection reason. Fixed budgets produce sparse coverage on long recordings; do not turn a sampled absence into “never appears.”
Start with a bounded scan. Identify relevant speech cues (“look here,” “this diagram”) and visual candidates, then inspect a tighter interval or explicit timestamps. A visual event need not be mentioned in speech; do not use captions as the only search index.
uv run --no-project --with youtube-transcript-api --with yt-dlp python "$SKILL_DIR/scripts/watch.py" "URL_OR_LOCAL_FILE" --engine local --start 02:00 --end 02:25 --detail transcript --timestamps 02:13 --max-frames 4 --no-dedup --question "Read the tool name and explain the demonstration"
Focus times and cues are absolute source times; seconds, MM:SS, and HH:MM:SS are accepted. Frames lie inside `[start, end)`. Requested cue time and actual decoded time are distinct. Cue frames reserve space in the total cap; an excessiv
Curated, production-grade skills, agents, hooks, rules, commands, utilities, and presets for AI coding agents. No magic, no demos — battle-tested workflows built for developers who use AI seriously.
Repo: Mathews-Tom/armory
Generates Architecture Decision Records capturing context, rationale, alternatives, and…
Build AI agents and automate Claude Code programmatically via the Claude Agent SDK and…
Audits and enhances FastAPI and REST API documentation: missing descriptions, response codes,…
Generate architecture diagrams as fully editable SVG with native AWS, Azure, and GCP icons…
Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness,…
Optimize and prepare figures for arXiv submission: format conversion (EPS/PDF/PNG/JPG), size…