Skip to content

/picking-a-format

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

From plugin
kreuzberg-dev-kreuzberg
9.3k7 skills1 MCP
Install
$ npx -y skills add kreuzberg-dev/kreuzberg --skill picking-a-format --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/picking-a-format

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

SKILL.md

picking-a-format.SKILL.md
name: picking-a-format
description: Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:bcdbffe958890e1c589f3ea44540c67c86540196a844e6de5c011faccdd345fc Source-Hash: blake3:5f5a598889c48eafed824c3d233459b3a360278d1175be7d6c4267d3aaf96d87 Schema-Version: v1 -->

Picking a format

Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.

| Knob | What it controls | Values | Default | | ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- | | `--format` | How the CLI prints the result | `text`, `json`, `toon` | `text` (`extract`), `json` (`batch`) | | `--content-format` | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html`, `json`, `doctags` | `plain` | | `--token-reduction` | Strip whitespace / boilerplate for LLM contexts | `off`, `light`, `moderate`, `aggressive`, `maximum` | `off` |

`--format json` returns an envelope wrapping the `ExtractedDocument` — the document lives under `.result` for `extract` and under `.results[]` for `batch`, with `content`, `metadata`, `tables`, and `images` as fields of that nested document. `--format text` prints just `content`. `--content-format` is what shows up inside that `content` field.

Decision tree

Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)

Examples

Feed a PDF directly into an LLM:

xberg extract paper.pdf --content-format markdown

Index a corpus into a RAG store with tables and headings preserved:

xberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.results[] | {content: .content, tables: .tables}'

Strip a file to bare text for a token-tight summarizer:

xberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate

Pull metadata only, ignore content:

xberg extract file.pdf --format json | jq '.result.metadata'

When in doubt

  • **Default to `markdown`** as the content format. It is the best

compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.

  • Reach for `plain` only when downstream cannot tolerate any markup.
  • Reach for `djot` only if you're already in a djot/pandoc pipeline.
  • Reach for `html` only when re-rendering for the web.
  • Reach for `json` for a heading-driven content tree, or `doctags` for Docling-compatible output.

Token-reduction (orthogonal)

`--token-reduction` collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any `--content-format`:

  • `off` (default), `light`, `moderate`, `aggressive`, `maximum`.

Use `moderate` as a safe starting point for LLM context windows. `maximum` is lossy — verify before relying on it.

See `references/cli-reference.md` for the full flag set and `references/configuration.md` for the equivalent `output_format` and `token_reduction` keys in `xberg.toml`.

Read more
Ships withkreuzberg-dev-kreuzberg

The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.

Get the whole plugin, auto-invoked
Stats
9,305
Stars
584
Forks
Active
Maintenance
Rust
Language
MIT
License
15h ago
Last commit
1y ago
Created

Repo: kreuzberg-dev/kreuzberg

Other skills on kreuzberg-dev-kreuzberg.