Skip to content
AI & Agents
Skill

/yao-ocr

OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, bank cards, business licenses, tables, handwritten documents, or any visual text content.

BOOST
From plugin
yao
8.1k12 skills1 agent
Install
$ npx -y skills add YaoApp/yao --skill yao-ocr --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/yao-ocr

Context preview

The summary Claude sees to decide when to auto-load this skill.

OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, bank cards, business licenses, tables, handwritten documents, or any visual text content.

SKILL.md

yao-ocr.SKILL.md
name: yao-ocr
description: OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, bank cards, business licenses, tables, handwritten documents, or any visual text content.

OCR Tools

Two tools for optical character recognition, supporting both VLM-OCR (vision language models) and traditional OCR APIs (Baidu, Google, Azure, PaddleOCR).

ocr_recognize

Extract text from images or PDF files using OCR.

Basic usage (plain text output):

tai tool ocr_recognize --source /path/to/image.png

With URL:

tai tool ocr_recognize --source https://example.com/document.jpg

Table extraction as Markdown:

tai tool ocr_recognize --source /path/to/table.png --type table --output_format markdown

Invoice structured extraction:

tai tool ocr_recognize --source /path/to/invoice.pdf --type invoice --output_format json

With specific provider:

tai tool ocr_recognize --source /path/to/doc.png --provider baidu

VLM-OCR with custom prompt:

tai tool ocr_recognize --source /path/to/doc.png --provider llm:qwen-ocr --prompt "只提取表格中的金额列"

PDF page range:

tai tool ocr_recognize --source /path/to/report.pdf --pages "1-5" --output_format markdown

| Parameter | Type | Required | Description | | ------------- | ------ | -------- | -------------------------------------------------------------------------------------------- | | source | string | yes | Image or PDF file path/URL to recognize | | provider | string | no | LLM connector ID (`llm:xxx`) or OCR settings key (`baidu`/`paddleocr`/`google`/`azure`). Auto-selects if omitted | | type | string | no | Recognition type (default: `general`). See type table below | | output_format | string | no | `text` (default), `json` (with coordinates/fields), or `markdown` (structured) | | mode | string | no | `accurate` (default, best quality) or `standard` (faster) | | language | string | no | Language hint (ISO 639-1, e.g. `en`, `zh`, `ja`). Auto-detected if omitted | | prompt | string | no | Custom instruction for VLM-OCR only, appended to system prompt. Ignored by traditional OCR | | pages | string | no | PDF page range, e.g. `1-5` or `1,3,7`. All pages if omitted | | extra | JSON | no | Provider-specific parameters as a JSON object |

Recognition types

| Type | Description | Best output_format | | ---------------- | -------------------------- | ------------------ | | `general` | General text (default) | text | | `table` | Table extraction | markdown | | `handwriting` | Handwritten text | text | | `document` | Document layout parsing | markdown | | `invoice` | Invoice (VAT) | json | | `receipt` | Receipt / ticket | json | | `id_card` | ID card | json | | `bank_card` | Bank card | json | | `license` | Business license | json | | `vehicle_license` | Vehicle license | json | | `passport` | Passport | json | | `license_plate` | License plate | json |

If the chosen provider does not support the requested type, it automatically degrades to `general` and annotates the response metadata with `degraded_from`. VLM-OCR supports all types via prompt adaptation.

ocr_providers

List available OCR providers and their supported recognition types.

tai tool ocr_providers

Returns a list of providers including VLM-OCR models (from LLM connectors with `ocr` capability) and traditional API providers (from OCR settings). Each entry includes `id`, `name`, `type` (`vlm` or `traditional`), and `supported_types`.

PDF support

| Provider | PDF | Notes | | ---------- | --- | ---------------------------------------- | | Baidu | yes | pdf_file parameter | | Azure | yes | Document Intelligence native support | | PaddleOCR | yes | pdf + fileType=0 | | Google | no | Sync API does not support PDF | | VLM (llm:) | no | Vision models accept images only |

For providers that do not support PDF, use Baidu, Azure, or PaddleOCR instead.

Multi-page PDF response

Multi-page PDFs are automatically split page-by-page. Instead of printing all text, the tool returns a JSON summary with file paths for each page result:

{
  "source": "report.pdf",
  "total_pages": 10,
  "pages": 3,
  "results": [
    {"page": 1, "file": ".tool-tmp/ocr-a1b2c3d4/page-1.txt", "preview": "Invoice No: INV-001..."},
    {"page": 2, "file": ".tool-tmp/ocr-a1b2c3d4/page-2.txt", "preview": "Invoice No: INV-002..."},
    {"page": 3, "file": ".tool-tmp/ocr-a1b2c3d4/page-3.txt", "preview": "Invoice No: INV-003..."}
  ]
}

To read full content of a specific page, use `cat`:

cat .tool-tmp/ocr-a1b2c3d4/page-2.txt

Single-page PDFs and images return inline text as usual (no file indirection).

Guidelines

  • Use `output_format=text` (default) when you just need the text content — simplest for LLM processing
  • Use `output_format=json` for structured types (invoice, id_card, etc.) to get key-value fields
  • Use `output_format=markdown` for documents and tables to preser
Read more
Ships withyao

✨ All your agents and workspaces in one place, on every device you own. Track tasks on a board, accessible from desktop, mobile, browser, or API. Self-hosted.

Get the whole plugin
Stats
8,072
Stars
717
Forks
Active
Maintenance
Go
Language
5d ago
Last commit
5y ago
Created
3h ago
Added

Repo: YaoApp/yao

Other skills on yao.