batch-extraction
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command,…
Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
$ npx -y skills add xberg-io/xberg --skill picking-a-format --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/picking-a-formatContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
name: picking-a-format description: Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:bcdbffe958890e1c589f3ea44540c67c86540196a844e6de5c011faccdd345fc Source-Hash: blake3:5f5a598889c48eafed824c3d233459b3a360278d1175be7d6c4267d3aaf96d87 Schema-Version: v1 -->
Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.
| Knob | What it controls | Values | Default | | ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- | | `--format` | How the CLI prints the result | `text`, `json`, `toon` | `text` (`extract`), `json` (`batch`) | | `--content-format` | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html`, `json`, `doctags` | `plain` | | `--token-reduction` | Strip whitespace / boilerplate for LLM contexts | `off`, `light`, `moderate`, `aggressive`, `maximum` | `off` |
`--format json` returns an envelope wrapping the `ExtractedDocument` — the document lives under `.result` for `extract` and under `.results[]` for `batch`, with `content`, `metadata`, `tables`, and `images` as fields of that nested document. `--format text` prints just `content`. `--content-format` is what shows up inside that `content` field.
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)Feed a PDF directly into an LLM:
xberg extract paper.pdf --content-format markdown
Index a corpus into a RAG store with tables and headings preserved:
xberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.results[] | {content: .content, tables: .tables}'Strip a file to bare text for a token-tight summarizer:
xberg extract long.pdf \ --content-format plain \ --token-reduction moderate
Pull metadata only, ignore content:
xberg extract file.pdf --format json | jq '.result.metadata'
compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.
`--token-reduction` collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any `--content-format`:
Use `moderate` as a safe starting point for LLM context windows. `maximum` is lossy — verify before relying on it.
See `references/cli-reference.md` for the full flag set and `references/configuration.md` for the equivalent `output_format` and `token_reduction` keys in `xberg.toml`.
The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.
Repo: xberg-io/xberg
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command,…
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers,…
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.…
Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown /…
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and…