/ai-provider-claude-vision
Image understanding and document analysis with Claude's multimodal capabilities -- image input formats, PDF processing, multi-image patterns, structured extraction, and token cost estimation
$ npx -y skills add agents-inc/skills --skill ai-provider-claude-vision --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
- You can call itInvoke it directly when you want it.
- Slash command
/ai-provider-claude-vision
Context preview
The summary Claude sees to decide when to auto-load this skill.
Image understanding and document analysis with Claude's multimodal capabilities -- image input formats, PDF processing, multi-image patterns, structured extraction, and token cost estimation
SKILL.md
ai-provider-claude-vision.SKILL.mdname: ai-provider-claude-vision
description: Image understanding and document analysis with Claude's multimodal capabilities -- image input formats, PDF processing, multi-image patterns, structured extraction, and token cost estimation
Claude Vision Patterns
> **Quick Guide:** Use `type: "image"` content blocks for images (base64, URL, or file_id) and `type: "document"` content blocks for PDFs. Supported image formats: JPEG, PNG, GIF, WebP. Images before text in the content array improves results. Token cost formula: `tokens = (width * height) / 750`. Images are auto-resized if the long edge exceeds 1568px or exceeds ~1600 tokens. PDFs use `type: "document"` with `media_type: "application/pdf"`. No OCR library needed -- Claude reads text directly from images and PDFs.
---
<critical_requirements>
CRITICAL: Before Using This Skill
> **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
**(You MUST use `type: "image"` for images and `type: "document"` for PDFs -- they are different content block types)**
**(You MUST place images and documents BEFORE text in the content array -- Claude performs better with visual content first)**
**(You MUST always provide `max_tokens` in every request -- it is required and has no default)**
**(You MUST iterate over `response.content` blocks -- never assume a single text block in the response)**
**(You MUST use named constants for max_tokens, token budgets, and pixel limits -- no magic numbers)**
</critical_requirements>
---
**Auto-detection:** Claude vision, image analysis, image input, base64 image, URL image, type image, type document, media_type image/jpeg, media_type image/png, image/webp, image/gif, application/pdf, PDF processing, document extraction, multimodal, multi-image, image comparison, chart analysis, screenshot analysis, image understanding, visual content, vision API
**When to use:**
- Sending images to Claude for analysis, description, or data extraction
- Processing PDF documents for text extraction, chart analysis, or summarization
- Comparing multiple images in a single request
- Extracting structured data from screenshots, receipts, charts, or forms
- Building document processing pipelines with Claude
- Estimating token costs for image-heavy workloads
**Key patterns covered:**
- Image input via base64, URL, and Files API
- PDF document input and processing
- Multi-image requests and comparison patterns
- Image + text prompting best practices
- Token cost estimation and image sizing
- Structured data extraction from visual content
- Multi-turn vision conversations
- Prompt caching with images and PDFs
**When NOT to use:**
- General Claude API usage without images or documents -- use the general Anthropic SDK patterns instead
- Image generation or editing -- Claude is understanding-only, it cannot create or modify images
- Identifying specific people in images -- Claude refuses to name people (Anthropic policy)
- Medical diagnostic imaging (CTs, MRIs) -- not designed for clinical diagnosis
---
Examples Index
- [Core: Image & PDF Input](examples/core.md) -- Base64, URL, file_id, PDF input, multi-image, token estimation
- [Extraction & Prompting](examples/extraction.md) -- Structured extraction, comparison, prompting best practices, caching
- [Quick API Reference](reference.md) -- Content block types, supported formats, size limits, token formula
---
<philosophy>
Philosophy
Claude's vision capabilities treat images and documents as **first-class content blocks** alongside text. There is no separate "vision API" -- you add image or document blocks to the same Messages API you already use for text.
**Core principles:**
1. **Images are content blocks, not attachments** -- Images and PDFs are content blocks in the `messages` array, interleaved with text. They are not uploaded separately or referenced by URL-only. 2. **Image-first ordering** -- Place images before text in the content array. This mirrors how `documents first, query last` improves text prompts. Claude processes visual content better when it sees the image before the question. 3. **No OCR needed** -- Claude reads text directly from images and PDFs. You do not need to pre-extract text with an OCR library. For PDFs, Claude processes both the extracted text and a rendered image of each page. 4. **Token costs scale with pixels** -- Image tokens are proportional to resolution: `tokens = (width * height) / 750`. Downsizing images before sending saves tokens without losing meaningful detail for most use cases. 5. **PDFs are dual-processed** -- Each PDF page is converted to an image AND has its text extracted. Claude sees both, giving it access to visual layout and textual content.
**When to use vision:**
- Analyzing screenshots, photos, charts, diagrams, or infographics
- Extracting data from forms, receipts, invoices, or tables
- Processing PDF documents for summarization, extraction, or analysis
- Comparing multiple images (before/after, A/B testing, design review)
- Understanding visual context that text alone cannot capture
**When NOT to use:**
- Pure text tasks with no visual component -- vision adds unnecessary token cost
- Tasks requiring pixel-perfect spatial precision -- Claude's spatial reasoning is approximate
- Identifying specific people -- Claude refuses to name individuals (Anthropic policy)
- Replacing professional medical imaging analysis (CTs, MRIs, X-rays)
</philosophy>
---
<patterns>
Core Patterns
Pattern 1: Base64 Image Input
Read a local file, encode to base64, send as `type: "image"` content block. Image block before text block.
// Image block first, text prompt second, iterate response content blocks
content: [
{
type: "image",
source: { type: "base64", media_type: "image/png", data: imageData },
},
{ type: "text", text: "Describe what you see in this image." },
];**Why good:** Image bef
Read more
name: ai-provider-claude-vision description: Image understanding and document analysis with Claude's multimodal capabilities -- image input formats, PDF processing, multi-image patterns, structured extraction, and token cost estimation
Claude Vision Patterns
> **Quick Guide:** Use `type: "image"` content blocks for images (base64, URL, or file_id) and `type: "document"` content blocks for PDFs. Supported image formats: JPEG, PNG, GIF, WebP. Images before text in the content array improves results. Token cost formula: `tokens = (width * height) / 750`. Images are auto-resized if the long edge exceeds 1568px or exceeds ~1600 tokens. PDFs use `type: "document"` with `media_type: "application/pdf"`. No OCR library needed -- Claude reads text directly from images and PDFs.
---
<critical_requirements>
CRITICAL: Before Using This Skill
> **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
**(You MUST use `type: "image"` for images and `type: "document"` for PDFs -- they are different content block types)**
**(You MUST place images and documents BEFORE text in the content array -- Claude performs better with visual content first)**
**(You MUST always provide `max_tokens` in every request -- it is required and has no default)**
**(You MUST iterate over `response.content` blocks -- never assume a single text block in the response)**
**(You MUST use named constants for max_tokens, token budgets, and pixel limits -- no magic numbers)**
</critical_requirements>
---
**Auto-detection:** Claude vision, image analysis, image input, base64 image, URL image, type image, type document, media_type image/jpeg, media_type image/png, image/webp, image/gif, application/pdf, PDF processing, document extraction, multimodal, multi-image, image comparison, chart analysis, screenshot analysis, image understanding, visual content, vision API
**When to use:**
- Sending images to Claude for analysis, description, or data extraction
- Processing PDF documents for text extraction, chart analysis, or summarization
- Comparing multiple images in a single request
- Extracting structured data from screenshots, receipts, charts, or forms
- Building document processing pipelines with Claude
- Estimating token costs for image-heavy workloads
**Key patterns covered:**
- Image input via base64, URL, and Files API
- PDF document input and processing
- Multi-image requests and comparison patterns
- Image + text prompting best practices
- Token cost estimation and image sizing
- Structured data extraction from visual content
- Multi-turn vision conversations
- Prompt caching with images and PDFs
**When NOT to use:**
- General Claude API usage without images or documents -- use the general Anthropic SDK patterns instead
- Image generation or editing -- Claude is understanding-only, it cannot create or modify images
- Identifying specific people in images -- Claude refuses to name people (Anthropic policy)
- Medical diagnostic imaging (CTs, MRIs) -- not designed for clinical diagnosis
---
Examples Index
- [Core: Image & PDF Input](examples/core.md) -- Base64, URL, file_id, PDF input, multi-image, token estimation
- [Extraction & Prompting](examples/extraction.md) -- Structured extraction, comparison, prompting best practices, caching
- [Quick API Reference](reference.md) -- Content block types, supported formats, size limits, token formula
---
<philosophy>
Philosophy
Claude's vision capabilities treat images and documents as **first-class content blocks** alongside text. There is no separate "vision API" -- you add image or document blocks to the same Messages API you already use for text.
**Core principles:**
1. **Images are content blocks, not attachments** -- Images and PDFs are content blocks in the `messages` array, interleaved with text. They are not uploaded separately or referenced by URL-only. 2. **Image-first ordering** -- Place images before text in the content array. This mirrors how `documents first, query last` improves text prompts. Claude processes visual content better when it sees the image before the question. 3. **No OCR needed** -- Claude reads text directly from images and PDFs. You do not need to pre-extract text with an OCR library. For PDFs, Claude processes both the extracted text and a rendered image of each page. 4. **Token costs scale with pixels** -- Image tokens are proportional to resolution: `tokens = (width * height) / 750`. Downsizing images before sending saves tokens without losing meaningful detail for most use cases. 5. **PDFs are dual-processed** -- Each PDF page is converted to an image AND has its text extracted. Claude sees both, giving it access to visual layout and textual content.
**When to use vision:**
- Analyzing screenshots, photos, charts, diagrams, or infographics
- Extracting data from forms, receipts, invoices, or tables
- Processing PDF documents for summarization, extraction, or analysis
- Comparing multiple images (before/after, A/B testing, design review)
- Understanding visual context that text alone cannot capture
**When NOT to use:**
- Pure text tasks with no visual component -- vision adds unnecessary token cost
- Tasks requiring pixel-perfect spatial precision -- Claude's spatial reasoning is approximate
- Identifying specific people -- Claude refuses to name individuals (Anthropic policy)
- Replacing professional medical imaging analysis (CTs, MRIs, X-rays)
</philosophy>
---
<patterns>
Core Patterns
Pattern 1: Base64 Image Input
Read a local file, encode to base64, send as `type: "image"` content block. Image block before text block.
// Image block first, text prompt second, iterate response content blocks
content: [
{
type: "image",
source: { type: "base64", media_type: "image/png", data: imageData },
},
{ type: "text", text: "Describe what you see in this image." },
];**Why good:** Image bef
Showing the first part of this file.
The official skills marketplace for Agents Inc. 150+ skills covering everything from React and Prisma to Redis, ElevenLabs, and infrastructure tooling. Pick the skills that match your stack and install them via Claude Code. Need more control?
Repo: agents-inc/skills
Other skills on agents-inc-skills.
- /ai-infrastructure-huggingface-inference
Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints
Open skill - /ai-infrastructure-litellm
LiteLLM proxy server setup, TypeScript client patterns via OpenAI SDK, model routing, fallbacks, load balancing, spend tracking, virtual keys, and production deployment
Open skill - /ai-infrastructure-modal
Serverless GPU compute platform for AI model deployment — web endpoints, GPU functions, model serving, and TypeScript client patterns
Open skill - /ai-infrastructure-ollama
Local LLM inference with the Ollama JavaScript client -- chat, streaming, tool calling, vision, embeddings, structured output, model management, and OpenAI-compatible endpoint
Open skill - /ai-infrastructure-replicate
Replicate SDK patterns for TypeScript/Node.js -- client setup, predictions, streaming, webhooks, file handling, model versioning, deployments, and training
Open skill - /ai-infrastructure-together-ai
Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation, fine-tuning, and OpenAI-compatible endpoints
Open skill

