/picking-a-format
Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
$ npx -y skills add kreuzberg-dev/kreuzberg --skill picking-a-format --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/picking-a-format
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
SKILL.md
picking-a-format.SKILL.mdname: picking-a-format
description: Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:470e563273f21e16138f65bd8a6c9b7a7b0c4adb003af8747c2957f3c17df0c4 Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->
Picking a format
Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.
| Knob | What it controls | Values | Default | | ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- | | `--format` | How the CLI prints the result | `text`, `json`, `toon` | `text` (`extract`), `json` (`batch`) | | `--content-format` | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html`, `json` | `plain` | | `--token-reduction` | Strip whitespace / boilerplate for LLM contexts | `off`, `light`, `moderate`, `aggressive`, `maximum` | `off` |
`--format json` returns an envelope wrapping the `ExtractedDocument` — the document lives under `.result` for `extract` and under `.results[]` for `batch`, with `content`, `metadata`, `tables`, and `images` as fields of that nested document. `--format text` prints just `content`. `--content-format` is what shows up inside that `content` field.
Decision tree
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)Examples
Feed a PDF directly into an LLM:
xberg extract paper.pdf --content-format markdown
Index a corpus into a RAG store with tables and headings preserved:
xberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.results[] | {content: .content, tables: .tables}'Strip a file to bare text for a token-tight summarizer:
xberg extract long.pdf \
--content-format plain \
--token-reduction moderate
Pull metadata only, ignore content:
xberg extract file.pdf --format json | jq '.result.metadata'
When in doubt
- **Default to `markdown`** as the content format. It is the best
compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.
- Reach for `plain` only when downstream cannot tolerate any markup.
- Reach for `djot` only if you're already in a djot/pandoc pipeline.
- Reach for `html` only when re-rendering for the web.
Token-reduction (orthogonal)
`--token-reduction` collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any `--content-format`:
- `off` (default), `light`, `moderate`, `aggressive`, `maximum`.
Use `moderate` as a safe starting point for LLM context windows. `maximum` is lossy — verify before relying on it.
See `references/cli-reference.md` for the full flag set and `references/configuration.md` for the equivalent `output_format` and `token_reduction` keys in `xberg.toml`.
Read more
name: picking-a-format description: Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:470e563273f21e16138f65bd8a6c9b7a7b0c4adb003af8747c2957f3c17df0c4 Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->
Picking a format
Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.
| Knob | What it controls | Values | Default | | ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- | | `--format` | How the CLI prints the result | `text`, `json`, `toon` | `text` (`extract`), `json` (`batch`) | | `--content-format` | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html`, `json` | `plain` | | `--token-reduction` | Strip whitespace / boilerplate for LLM contexts | `off`, `light`, `moderate`, `aggressive`, `maximum` | `off` |
`--format json` returns an envelope wrapping the `ExtractedDocument` — the document lives under `.result` for `extract` and under `.results[]` for `batch`, with `content`, `metadata`, `tables`, and `images` as fields of that nested document. `--format text` prints just `content`. `--content-format` is what shows up inside that `content` field.
Decision tree
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)Examples
Feed a PDF directly into an LLM:
xberg extract paper.pdf --content-format markdown
Index a corpus into a RAG store with tables and headings preserved:
xberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.results[] | {content: .content, tables: .tables}'Strip a file to bare text for a token-tight summarizer:
xberg extract long.pdf \ --content-format plain \ --token-reduction moderate
Pull metadata only, ignore content:
xberg extract file.pdf --format json | jq '.result.metadata'
When in doubt
- **Default to `markdown`** as the content format. It is the best
compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.
- Reach for `plain` only when downstream cannot tolerate any markup.
- Reach for `djot` only if you're already in a djot/pandoc pipeline.
- Reach for `html` only when re-rendering for the web.
Token-reduction (orthogonal)
`--token-reduction` collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any `--content-format`:
- `off` (default), `light`, `moderate`, `aggressive`, `maximum`.
Use `moderate` as a safe starting point for LLM context windows. `maximum` is lossy — verify before relying on it.
See `references/cli-reference.md` for the full flag set and `references/configuration.md` for the equivalent `output_format` and `token_reduction` keys in `xberg.toml`.
The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.
Repo: kreuzberg-dev/kreuzberg
Other skills on xberg.
- /batch-extraction
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.
Open skill - /chunking
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.
Open skill - /extracting-keywords
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.
Open skill - /extracting-tables
Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown / JSON cells), and known limits.
Open skill - /extracting-with-ocr
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
Open skill - /xberg
Extract text, tables, metadata, and images from 101 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg. Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async),
Open skill

