Skip to content
Data
Skill

/extracting-keywords

Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

From plugin
xberg
8.9k7 skills1 MCP
Install
$ npx -y skills add kreuzberg-dev/kreuzberg --skill extracting-keywords --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/extracting-keywords

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

SKILL.md

extracting-keywords.SKILL.md
name: extracting-keywords
description: Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0da13ce50f1fef8c192e912e7c2b0d039209a7b676829abe13f4fc86b5ea20bc Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->

Extracting keywords, language, and embeddings

Use this for the enrichment surface around extraction: statistical keyword extraction, language detection, and vector embeddings. Keywords and language detection ride along with extraction and land on the result; embeddings are produced by a dedicated `embed` command.

Keywords (YAKE / RAKE)

Keyword extraction is configured via the `[keywords]` config block (or inline JSON) — there is no single `--keywords` CLI flag. When enabled, extracted keywords appear on `result.extracted_keywords` (`extractedKeywords` in Node.js; the CLI JSON field is `extracted_keywords`). Two algorithms are available:

  • **YAKE** (`"yake"`) — statistical, unsupervised single-document

extraction. Good general default.

  • **RAKE** (`"rake"`) — co-occurrence / phrase-based. Favors multi-word

key phrases.

> Feature-gated: keyword extraction requires the CLI to be built with the > `keywords-yake` and/or `keywords-rake` Cargo features (both are in the > default/`full` build). If the CLI was built without them, the `[keywords]` > config block is silently ignored — `result.extracted_keywords` simply stays empty > rather than erroring. The `"yake"` algorithm needs `keywords-yake`; `"rake"` > needs `keywords-rake`.

Enable via inline JSON on the CLI:

xberg extract paper.pdf --format json \
  --config-json '{"keywords":{"algorithm":"yake","max_keywords":15,"language":"en"}}' \
  | jq '.extracted_keywords'

Or in a config file:

[keywords]
algorithm = "rake"       # "yake" or "rake"
max_keywords = 10        # default 10
min_score = 0.0          # filter below this score (normalized 0.0-1.0 for both algorithms)
ngram_range = [1, 3]     # unigrams..trigrams (default); config-file only
language = "en"          # stopword language; omit to skip stopword filtering
xberg extract report.pdf --config xberg.toml --format json | jq '.extracted_keywords'

Field notes:

  • `max_keywords` caps how many keywords are returned (default 10).
  • `min_score` filters low-scoring keywords. Both YAKE and RAKE normalize

their scores to the `0.0`-`1.0` range with *higher-is-better*, so `min_score` retains keywords with `score >= min_score` identically for either algorithm.

  • `ngram_range` is `[min, max]`: `[1,1]` unigrams only, `[1,2]` adds

bigrams, `[1,3]` (default) adds trigrams. Config-file only — it is not a field on the language bindings' `KeywordConfig`.

  • `language` enables stopword filtering for that language; omit it to

disable stopword filtering entirely.

Language detection

Language detection is a real CLI flag: `--detect-language`. Detected languages appear on `result.detected_languages`:

xberg extract multilingual.pdf --detect-language true --format json \
  | jq '.detected_languages'

In a config file it lives under `[language_detection]`:

[language_detection]
enabled = true
min_confidence = 0.8
detect_multiple = false

The CLI flag enables detection with `min_confidence = 0.8` and single-language mode; use the config block to detect multiple languages or tune confidence.

Embeddings (`embed` command)

The standalone `embed` command produces vector embeddings for text from `--text` (repeatable) or stdin. It does not run extraction — pipe extracted content in if you want document embeddings.

# Local ONNX preset model (default provider)
xberg embed --text "first passage" --text "second passage" --preset balanced

# Embed extracted document text
xberg extract report.pdf | xberg embed --preset quality

Presets for the local provider: `fast`, `balanced` (default), `quality`, `multilingual`. Output defaults to JSON (`--format json`).

`--provider` selects the embedding source:

| Provider | Flag | Notes | | -------- | ------------------------------------- | --------------------------------------------- | | `local` | `--preset <fast\|balanced\|quality\|multilingual>` | **Default.** ONNX model, no API key. | | `llm` | `--model <id>` `--api-key <key>` | liter-llm routing, e.g. `openai/text-embedding-3-small`. | | `plugin` | `--plugin <name>` | A backend pre-registered in-process via the plugin API. |

# Provider-hosted embeddings via an LLM
xberg embed --text "query text" \
  --provider llm --model openai/text-embedding-3-small --api-key "$OPENAI_API_KEY"

Local embedding presets must be downloaded first if not cached. Pre-warm them with the cache command:

xberg cache warm --embedding-model balanced   # one preset
xberg cache warm --all-embeddings             # all available presets (currently 8)

Programmatic access

Keywords and detected languages live on the document in the result envelope:

from xberg import ExtractInput, extract, ExtractionConfig, KeywordConfig, KeywordAlgorithm

config = ExtractionConfig(
    keywords=KeywordConfig(algorithm=KeywordAlgorithm.YAKE, max_keywords=15, language="en"),
)
result = await extract(ExtractInput(uri="paper.pdf"), config)
doc = result.results[0]
print(doc.extracted_keywords)   # extracted keywords (when enabled)
print(doc.detected_languages)   # detected languages (when enabled)

See `references/python-api.md` and `references/configuration.md` in the sibling `xberg` skill for the keyword / language-detection config classes and the embedding presets.

##

Read more
Ships withxberg

The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.

Get the whole plugin
Stats
8,946
Stars
540
Forks
Active
Maintenance
Rust
Language
MIT
License
3h ago
Last commit
1y ago
Created

Repo: kreuzberg-dev/kreuzberg