chunking
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers,…
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.
$ npx -y skills add xberg-io/xberg --skill batch-extraction --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/batch-extractionContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.
name: batch-extraction description: Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:3a70b47dd77fcd3f37e83bed3d32b1ad1ea9d6c1642ae8e0ff7b09b55939858a Source-Hash: blake3:5f5a598889c48eafed824c3d233459b3a360278d1175be7d6c4267d3aaf96d87 Schema-Version: v1 -->
Use this when processing a directory or glob of documents in one pass. `xberg batch` shares one extraction config across every file, runs extractions concurrently, and returns one structured array — failures on individual files do not abort the run.
# Glob expands to many paths; results come back as a JSON array (default) xberg batch *.pdf # Mixed formats, markdown content for LLM ingestion xberg batch docs/*.docx --content-format markdown # Recurse with the shell, then extract xberg batch $(find ./corpus -name '*.pdf')
`batch` defaults to `--format json` (vs `--format text` for single `extract`). Each array entry is a full extraction result, so downstream code can index by position into the input path list.
xberg batch reports/*.pdf \
| jq '.[] | {chars: (.content | length), mime: .mime_type}'`--max-concurrent` caps how many files extract at once. When omitted, the scheduler derives document concurrency from the total thread budget. Lower it on memory-constrained hosts or when OCR/ML models are active, since each in-flight extraction holds its own buffers. Layout-heavy batches are further limited (1 concurrent extraction for all-PDF-layout batches, 2 for mixed layout):
# Cap at 4 concurrent extractions xberg batch scans/*.pdf --ocr true --max-concurrent 4
`--max-threads` additionally caps *total* internal threads (Rayon, ONNX intra-op, the batch semaphore) for tightly constrained environments:
xberg batch *.pdf --max-concurrent 2 --max-threads 4
A single shared config does not always fit. `--file-configs` points at a JSON file mapping each path to its own override object, merged on top of the shared config for that file only:
{
"scan.pdf": { "force_ocr": true },
"report.pdf": { "output_format": "markdown" },
"data.xlsx": { "output_format": "json" }
}xberg batch scan.pdf report.pdf data.xlsx --file-configs overrides.json
Keys are file paths (matching the paths passed on the command line); values are per-file extraction config objects in snake_case, the same shape as a config file.
For text/toon output with image extraction, `--output-dir` controls where referenced image files (e.g. `image_0.png`) are written; the directory must already exist. JSON output embeds image bytes inline and ignores `--output-dir`.
mkdir -p out/images xberg batch slides/*.pptx --extract-images true --output-dir out/images --format text
Batch extraction is fault-tolerant per file: one unreadable or corrupt document does not stop the rest. Inspect results for partial content and surfaced errors rather than relying on the process exit code alone. Pair with `--max-concurrent` to avoid exhausting memory when a few large files sit in a big batch.
Every `extract` flag also applies to `batch` (OCR, chunking, layout, content format, etc.) and is shared across all files unless a `--file-configs` entry overrides it:
xberg batch invoices/*.pdf \ --layout --layout-table-model slanet_wireless \ --content-format markdown --max-concurrent 8
A config file works too and auto-discovers from the cwd upward:
output_format = "markdown" [ocr] backend = "tesseract" language = "eng"
xberg batch corpus/*.pdf --config xberg.toml
From Python, `extract_batch` takes a list of `ExtractInput`s and returns one envelope whose `results` array holds a document per input:
from xberg import ExtractInput, extract_batch, ExtractionConfig
config = ExtractionConfig(output_format="markdown")
inputs = [ExtractInput(uri=p) for p in ["a.pdf", "b.docx", "c.xlsx"]]
output = await extract_batch(inputs, config)
for doc in output.results:
print(len(doc.content))Per-input overrides go on `ExtractInput.config` (a `FileExtractionConfig`). Node.js mirrors this with `extractBatch`; Rust uses `extract_batch(inputs, &config)`. See `references/python-api.md`, `references/nodejs-api.md`, and `references/rust-api.md` in the sibling `xberg` skill.
When the `xberg` MCP server is registered, prefer the `extract_batch` tool over shelling out — it takes an array of input objects and a config object and returns structured results directly.
`extract` to `--format text`. Set `--format` explicitly if a script depends on one shape.
`--max-concurrent` ceiling.
command line, not absolute-resolved variants.
See `references/cli-reference.md` for the full `batch` flag set.
The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.
Repo: xberg-io/xberg
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers,…
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.…
Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown /…
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and…
Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the…