Skip to content
Data
Skill

/extracting-keywords

Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

From plugin
xberg
8.9k7 skills1 MCP
Install
$ npx -y skills add xberg-io/xberg --skill extracting-keywords --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/extracting-keywords

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

SKILL.md

extracting-keywords.SKILL.md
name: extracting-keywords
description: Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0da13ce50f1fef8c192e912e7c2b0d039209a7b676829abe13f4fc86b5ea20bc Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->

Extracting keywords, language, and embeddings

Use this for the enrichment surface around extraction: statistical keyword extraction, language detection, and vector embeddings. Keywords and language detection ride along with extraction and land on the result; embeddings are produced by a dedicated `embed` command.

Keywords (YAKE / RAKE)

Keyword extraction is configured via the `[keywords]` config block (or inline JSON) — there is no single `--keywords` CLI flag. When enabled, extracted keywords appear on `result.extracted_keywords` (`extractedKeywords` in Node.js; the CLI JSON field is `extracted_keywords`). Two algorithms are available:

  • **YAKE** (`"yake"`) — statistical, unsupervised single-document

extraction. Good general default.

  • **RAKE** (`"rake"`) — co-occurrence / phrase-based. Favors multi-word

key phrases.

> Feature-gated: keyword extraction requires the CLI to be built with the > `keywords-yake` and/or `keywords-rake` Cargo features (both are in the > default/`full` build). If the CLI was built without them, the `[keywords]` > config block is silently ignored — `result.extracted_keywords` simply stays empty > rather than erroring. The `"yake"` algorithm needs `keywords-yake`; `"rake"` > needs `keywords-rake`.

Enable via inline JSON on the CLI:

xberg extract paper.pdf --format json \
  --config-json '{"keywords":{"algorithm":"yake","max_keywords":15,"language":"en"}}' \
  | jq '.extracted_keywords'

Or in a config file:

[keywords]
algorithm = "rake"       # "yake" or "rake"
max_keywords = 10        # default 10
min_score = 0.0          # filter below this score (normalized 0.0-1.0 for both algorithms)
ngram_range = [1, 3]     # unigrams..trigrams (default); config-file only
language = "en"          # stopword language; omit to skip stopword filtering
xberg extract report.pdf --config xberg.toml --format json | jq '.extracted_keywords'

Field notes:

  • `max_keywords` caps how many keywords are returned (default 10).
  • `min_score` filters low-scoring keywords. Both YAKE and RAKE normalize

their scores to the `0.0`-`1.0` range with *higher-is-better*, so `min_score` retains keywords with `score >= min_score` identically for either algorithm.

  • `ngram_range` is `[min, max]`: `[1,1]` unigrams only, `[1,2]` adds

bigrams, `[1,3]` (default) adds trigrams. Config-file only — it is not a field on the language bindings' `KeywordConfig`.

  • `language` enables stopword filtering for that language; omit it to

disable stopword filtering entirely.

Language detection

Language detection is a real CLI flag: `--detect-language`. Detected languages appear on `result.detected_languages`:

xberg extract multilingual.pdf --detect-language true --format json \
  | jq '.detected_languages'

In a config file it lives under `[language_detection]`:

[language_detection]
enabled = true
min_confidence = 0.8
detect_multiple = false

The CLI flag enables detection with `min_confidence = 0.8` and single-language mode; use the config block to detect multiple languages or tune confidence.

Embeddings (`embed` command)

The standalone `embed` command produces vector embeddings for text from `--text` (repeatable) or stdin. It does not run extraction — pipe extracted content in if you want document embeddings.

# Local ONNX preset model (default provider)
xberg embed --text "first passage" --text "second passage" --preset balanced

# Embed extracted document text
xberg extract report.pdf | xberg embed --preset quality

Presets for the local provider: `fast`, `balanced` (default), `quality`, `multilingual`. Output defaults to JSON (`--format json`).

`--provider` selects the embedding source:

| Provider | Flag | Notes | | -------- | ------------------------------------- | --------------------------------------------- | | `local` | `--preset <fast\|balanced\|quality\|multilingual>` | **Default.** ONNX model, no API key. | | `llm` | `--model <id>` `--api-key <key>` | liter-llm routing, e.g. `openai/text-embedding-3-small`. | | `plugin` | `--plugin <name>` | A backend pre-registered in-process via the plugin API. |

# Provider-hosted embeddings via an LLM
xberg embed --text "query text" \
  --provider llm --model openai/text-embedding-3-small --api-key "$OPENAI_API_KEY"

Local embedding presets must be downloaded first if not cached. Pre-warm them with the cache command:

xberg cache warm --embedding-model balanced   # one preset
xberg cache warm --all-embeddings             # all available presets (currently 8)

Programmatic access

Keywords and detected languages live on the document in the result envelope:

from xberg import ExtractInput, extract, ExtractionConfig, KeywordConfig, KeywordAlgorithm

config = ExtractionConfig(
    keywords=KeywordConfig(algorithm=KeywordAlgorithm.YAKE, max_keywords=15, language="en"),
)
result = await extract(ExtractInput(uri="paper.pdf"), config)
doc = result.results[0]
print(doc.extracted_keywords)   # extracted keywords (when enabled)
print(doc.detected_languages)   # detected languages (when enabled)

See `references/python-api.md` and `references/configuration.md` in the sibling `xberg` skill for the keyword / language-detection config classes and the embedding presets.

##

Read more
Ships withxberg

The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.

Get the whole plugin