Skip to content
Data
Skill

/chunking

Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.

From plugin
xberg-io-xberg
9.3k7 skills1 MCP
Install
$ npx -y skills add xberg-io/xberg --skill chunking --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/chunking

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.

SKILL.md

chunking.SKILL.md
name: chunking
description: Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.

<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0b5cd4bec9d2a8f3452e07139ef56d6b87f9f7990714d15483a03435e751af59 Source-Hash: blake3:5f5a598889c48eafed824c3d233459b3a360278d1175be7d6c4267d3aaf96d87 Schema-Version: v1 -->

Chunking

Use this when feeding documents into an LLM context window or a vector store. Xberg chunks two ways: inline during extraction (chunks land on each document's `chunks` field), or standalone via the `chunk` command for text you already have. Sizing is character-based by default, or token-based when a tokenizer model is supplied.

Inline during extraction

Turn on chunking with `--chunk` and the chunks appear on the structured result under `chunks`:

# 1000-char chunks, 200-char overlap (defaults when --chunk is on)
xberg extract report.pdf --chunk --format json | jq '.chunks | length'

# Explicit size + overlap
xberg extract report.pdf --chunk --chunk-size 1500 --chunk-overlap 300 --format json

Overlap must be smaller than chunk size — the CLI rejects `--chunk-overlap >= --chunk-size`. When you set only `--chunk-overlap` against an existing config, an overlap that exceeds the size is clamped to `chunk_size / 4`.

Standalone `chunk` command

Chunk text you already have, from `--text` or stdin. Output defaults to JSON:

# From a flag
xberg chunk --text "long document text ..." --chunk-size 800 --chunk-overlap 100

# From stdin (pipe extracted content straight in)
xberg extract notes.md | xberg chunk --chunk-size 500 --format json

JSON output carries `chunks` (array of strings), `chunk_count`, the resolved `config` (`max_characters`, `overlap`, `chunker_type`), and `input_size_bytes`. Use `--format text` for a human-readable dump with `--- chunk N ---` separators.

> Note: in the JSON output, `chunker_type` is rendered capitalized (`"Text"`, > `"Markdown"`, `"Yaml"`, `"Semantic"`) because it is emitted via Rust's Debug > formatting, whereas the `--chunker-type` input flag is lowercase > (`text`, `markdown`, `yaml`, `semantic`). Lowercase the value before > comparing if you parse it back.

Chunker types

`--chunker-type` selects the splitting strategy (standalone `chunk` command):

| Type | Behavior | | ---------- | ------------------------------------------------------------------- | | `text` | **Default.** Plain character-window splitting with overlap. | | `markdown` | Markdown-aware — splits on structure (headings, blocks) where possible. | | `yaml` | YAML-aware splitting for structured config/data documents. | | `semantic` | Topic-boundary splitting driven by `--topic-threshold` (0.0–1.0, default 0.75). |

# Markdown-aware chunking keeps headings and blocks intact
xberg chunk --text "$(cat README.md)" --chunker-type markdown

# Semantic chunking — lower threshold = more, smaller topic chunks
xberg chunk --text "$(cat transcript.txt)" --chunker-type semantic --topic-threshold 0.6

Token-based sizing

By default `--chunk-size` counts characters. To size chunks by tokens for a specific model, pass `--chunking-tokenizer` with a HuggingFace tokenizer id. On the `extract` command this implicitly enables chunking. Requires the `chunking-tokenizers` feature (present in the default CLI build).

# Size chunks by GPT-4o tokens during extraction
xberg extract report.pdf --chunking-tokenizer Xenova/gpt-4o --format json

# Or on the standalone command
xberg chunk --text "$(cat doc.txt)" --chunking-tokenizer Xenova/gpt-4o --chunk-size 512

With a tokenizer set, `--chunk-size` is interpreted in tokens, not characters.

Config file alternative

Field names in config files are snake_case under `[chunking]`:

[chunking]
max_characters = 1000
overlap = 200
chunker_type = "markdown"
xberg extract report.pdf --config xberg.toml --format json

> CLI flags map to config fields as `--chunk-size` → `max_characters` and > `--chunk-overlap` → `overlap`. In config files use the snake_case names.

Programmatic access

From Python, enable chunking on the config and read the chunks off the document in the result envelope (`result.results[0].chunks`):

from xberg import ExtractInput, extract, ExtractionConfig, ChunkingConfig

config = ExtractionConfig(
    chunking=ChunkingConfig(max_characters=1000, overlap=200),
)
result = await extract(ExtractInput(uri="report.pdf"), config)
for chunk in result.results[0].chunks or []:
    print(len(chunk.content))

> The public Python `ChunkingConfig` (a dataclass) uses constructor kwargs > `max_characters` / `overlap`; the Rust core struct fields are also > `max_characters` / `overlap`. TOML/JSON config keys are `max_chars` / > `max_overlap` (with `max_characters` / `overlap` accepted as serde aliases), > and dict-form config passed to `ExtractionConfig` likewise accepts the > `max_chars` / `max_overlap` aliases; Node's `ChunkingConfig` interface uses > `maxCharacters` / `overlap`. See `references/python-api.md` and > `references/rust-api.md` in the sibling `xberg` skill.

Picking parameters

  • **RAG / vector store** — 500–1000 chars (or 256–512 tokens) with

10–20% overlap. Use `markdown` chunking for docs to keep sections whole.

  • **LLM summarization** — larger chunks (1500–4000 chars) with small

overlap; size by tokens to stay under the model window.

  • **Topic segmentation** — `semantic` chunker; tune `--topic-threshold`

down for finer splits, up for coarser ones.

Common pitfalls

  • **Overlap ≥ size** — rejected on `extract`; clamped to `size / 4` when

only overlap is changed against an existing config.

  • **Tokenizer without the feature** — `--chunking-tokenizer` errors if the

CLI w

Read more
Ships withxberg-io-xberg

The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.

Get the whole plugin
Stats
9,304
Stars
584
Forks
Active
Maintenance
Rust
Language
MIT
License
15h ago
Last commit
1y ago
Created

Repo: xberg-io/xberg

Other skills on xberg-io-xberg.