ork-assess
Assess a code change, design, architecture, workflow, or competing options against explicit criteria and evidence. Use when a request asks to assess, rate,…
Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via `gen4_turbo`), or building multimodal AI pipelines.
$ npx -y skills add yonatangross/orchestkit --skill multimodal-llm --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/multimodal-llmContext preview
The summary Claude sees to decide when to auto-load this skill.
Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via `gen4_turbo`), or building multimodal AI pipelines.
name: multimodal-llm license: MIT compatibility: "Claude Code 2.1.251+." author: OrchestKit version: 2.1.1 description: Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via `gen4_turbo`), or building multimodal AI pipelines. tags: [vision, audio, video, multimodal, image, speech, transcription, tts, kling, sora, veo, video-generation] user-invocable: false disable-model-invocation: true context: fork agent: multimodal-specialist complexity: high persuasion-type: reference effort: high metadata: category: mcp-enhancement allowed-tools: - Read - Glob - Grep - WebFetch - WebSearch
Integrate vision, audio, and video generation capabilities from leading multimodal models. Covers image analysis, document understanding, real-time voice agents, speech-to-text, text-to-speech, and AI video generation (Kling v3, Sora 2, Veo 3.1 std/lite/fast tiers, Runway Gen-4.5 via `gen4_turbo`).
> **Canonical model IDs** (pinned against `yonatan-hq/platform/apps/api/app/config.py`): > > | Provider | Model IDs | > |----------|-----------| > | Anthropic | `claude-opus-5` (recommended, 2,576 px budget, production default), `claude-opus-4-8`, `claude-opus-4-7`, `claude-opus-4-6`, `claude-sonnet-4-6`, `claude-haiku-4-5-20251001`. `claude-fable-5` is Anthropic's **frontier tier above Opus** (GA 2026-07). Premium cost — never auto-pin it; the fable-spend-consent gate requires explicit user consent before any Fable spend | > | OpenAI | `gpt-5.5` (current flagship) | > | Google | `gemini-3.1-pro-preview` (flagship), `gemini-3.1-flash-lite-preview` (cost) | > | Veo | `veo-3.1-generate-preview` / `veo-3.1-lite-generate-preview` / `veo-3.1-fast-generate-preview` | > | Kling | `kling-v3` (model_name field in Kling API) | > | Runway | `gen4_turbo` (product label: Gen-4.5) |
| Category | Rules | Impact | When to Use | |----------|-------|--------|-------------| | [Vision: Image Analysis](#vision-image-analysis) | 1 | HIGH | Image captioning, VQA, multi-image comparison, object detection | | [Vision: Document Understanding](#vision-document-understanding) | 1 | HIGH | OCR, chart/diagram analysis, PDF processing, table extraction | | [Vision: Model Selection](#vision-model-selection) | 1 | MEDIUM | Choosing provider, cost optimization, image size limits | | [Audio: Speech-to-Text](#audio-speech-to-text) | 1 | HIGH | Transcription, speaker diarization, long-form audio | | [Audio: Text-to-Speech](#audio-text-to-speech) | 1 | MEDIUM | Voice synthesis, expressive TTS, multi-speaker dialogue | | [Audio: Model Selection](#audio-model-selection) | 1 | MEDIUM | Real-time voice agents, provider comparison, pricing | | [Video: Model Selection](#video-model-selection) | 1 | HIGH | Choosing video gen provider (Kling, Sora, Veo, Runway) | | [Video: API Patterns](#video-api-patterns) | 1 | HIGH | Async task polling, SDK integration, webhook callbacks | | [Video: Multi-Shot](#video-multi-shot) | 1 | HIGH | Storyboarding, character elements, scene consistency |
**Total: 9 rules across 3 categories (Vision, Audio, Video Generation)**
Send images to multimodal LLMs for captioning, visual QA, and object detection. Always set `max_tokens` and resize images before encoding.
| Rule | File | Key Pattern | |------|------|-------------| | Image Analysis | `rules/vision-image-analysis.md` | Base64 encoding, multi-image, bounding boxes |
Extract structured data from documents, charts, and PDFs using vision models.
| Rule | File | Key Pattern | |------|------|-------------| | Document Vision | `rules/vision-document.md` | PDF page ranges, detail levels, OCR strategies |
Choose the right vision provider based on accuracy, cost, and context window needs.
| Rule | File | Key Pattern | |------|------|-------------| | Vision Models | `rules/vision-models.md` | Provider comparison, token costs, image limits |
Convert audio to text with speaker diarization, timestamps, and sentiment analysis.
| Rule | File | Key Pattern | |------|------|-------------| | Speech-to-Text | `rules/audio-speech-to-text.md` | Gemini long-form, GPT-4o-Transcribe, AssemblyAI features |
Generate natural speech from text with voice selection and expressive cues.
| Rule | File | Key Pattern | |------|------|-------------| | Text-to-Speech | `rules/audio-text-to-speech.md` | Gemini TTS, voice config, auditory cues |
Select the right audio/voice provider for real-time, transcription, or TTS use cases.
| Rule | File | Key Pattern | |------|------|-------------| | Audio Models | `rules/audio-models.md` | Real-time voice comparison, STT benchmarks, pricing |
Choose the right video generation provider based on use case, duration, and budget.
| Rule | File | Key Pattern | |------|------|-------------| | Video Models | `rules/video-generation-models.md` | Kling vs Sora vs Veo vs Runway, pricing, capabilities |
Integrate video generation APIs with proper async polling, SDKs, and webhook callbacks.
| Rule | File | Key Pattern | |------|------|-------------| | API Integration | `rules/video-generation-patterns.md` | Kling REST, fal.ai SDK, Vercel AI SDK, task polling |
Generate multi-scene videos with consistent characters using storyboarding and character elements.
| Rule | File | Key Pattern | |------|------|-------------| | Multi-Shot | `rules/video-multi-shot.md` | Kling v3 character elements, 6-shot storyboards, identity binding |
| Decision | Recommendation | |----------|----------------| | High accuracy vision | `claude-opus-5` (production default, 2,576 px vision b
The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.
Repo: yonatangross/orchestkit
Assess a code change, design, architecture, workflow, or competing options against explicit criteria and evidence. Use when a request asks to assess, rate,…
Compare plausible implementation, architecture, product, or operational approaches before committing to one. Use when a request asks to brainstorm, think…
Map an unfamiliar codebase, feature, architecture, data flow, or operational path with file-backed evidence. Use when a request asks how a system works, where…
Make an approved, scoped change and prove the affected behavior. Use when a request asks to implement, build, add, or land a feature that already has an agreed…
Review a pull request or branch for correctness, regressions, security, operational risk, and missing evidence. Use when a request asks to review a PR, review…
Verify that existing work is ready to merge, release, or hand off using an explicit evidence contract. Use when a request asks to verify, validate, prove,…