batch-extraction
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command,…
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
$ npx -y skills add xberg-io/xberg --skill extracting-with-ocr --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/extracting-with-ocrContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
name: extracting-with-ocr description: Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:3ec8b7cf60f56cbe5cc15a5a0b29d0c2f8e3cf4c3823503eb1128fb7f9ee11db Source-Hash: blake3:5f5a598889c48eafed824c3d233459b3a360278d1175be7d6c4267d3aaf96d87 Schema-Version: v1 -->
Use this when a document is image-based: scanned PDFs, photographed pages, screenshots, JPEG/PNG/TIFF with text. Xberg auto-OCRs raster images and auto-detects PDFs that lack a text layer. Force it on when extraction returned empty/garbled text from a PDF that "looks" textual.
xberg extract scan.pdf --force-ocr=true xberg extract scan.pdf --ocr=true --ocr-language eng
If a page has an unreliable text layer, `--force-ocr=true` re-rasterizes and runs OCR on every page.
Tesseract is the default and ships with the CLI — no extra install. Other backends are opt-in:
| Backend | Flag | Install | Notes | | ------------- | ------------------------------------- | ------------------------------------------------ | -------------------------------------------------------------- | | Tesseract | `--ocr-backend tesseract` (default) | bundled | Best general-purpose, 100+ languages via tessdata. | | PaddleOCR | `--ocr-backend paddle-ocr` | bundled (ONNX Runtime) | Strong on Asian scripts. Not available on WASM or Windows. | | Candle VLM | `--ocr-backend candle-trocr` (and other `candle-*`) | bundled (Candle) | Local vision OCR models (`candle-trocr`, `candle-paddleocr-vl`, `candle-glm-ocr`, `candle-deepseek-ocr`). | | VLM (hosted) | `--ocr-backend vlm` + `--vlm-model` | liter-llm provider (`--vlm-api-key`) | Multimodal LLM via liter-llm. Use when OCR fails on dense or handwritten layouts. |
Pick Tesseract first. Switch only when accuracy is unacceptable.
Tesseract uses ISO 639-2 codes. Default is `eng`. Combine with `+`:
xberg extract menu.jpg --ocr=true --ocr-language "eng+deu" xberg extract bilingual.pdf --ocr-language "eng+jpn" xberg extract any.pdf --ocr-language all # all installed packs
Install missing packs at the OS level:
# macOS brew install tesseract-lang # Debian/Ubuntu sudo apt install tesseract-ocr-deu tesseract-ocr-jpn tesseract-ocr-fra # Specific lang only sudo apt install tesseract-ocr-<iso639-2>
Xberg fails fast with a helpful error if you request a language pack that is not installed. Read the error — it names the missing file.
paddle-ocr / auto-rotate / layout models.
instant. Do not pass `--no-cache=true` unless you have a reason.
pool parallelizes across CPU cores. Cap with `--max-concurrent N` if memory is tight.
DPI is slower; 200 is usually enough for printed text.
classifier adds latency.
paddle-ocr and layout detection.
Long flag chains belong in `xberg.toml` — auto-discovered from cwd upward.
force_ocr = true output_format = "markdown" [ocr] backend = "tesseract" language = "eng+deu" auto_rotate = true
Then just run:
xberg extract document.pdf
bogus zero-width text layer. Re-run with `--force-ocr=true`.
passed via `--ocr-language`; consider `paddle-ocr` for Chinese/Japanese.
See `references/cli-reference.md` and `references/configuration.md` in the sibling `xberg` skill for the full flag and config schema.
The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.
Repo: xberg-io/xberg
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command,…
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers,…
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.…
Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown /…
Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the…