/extracting-keywords
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.
$ npx -y skills add xberg-io/xberg --skill extracting-keywords --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/extracting-keywords
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.
SKILL.md
extracting-keywords.SKILL.mdname: extracting-keywords
description: Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0da13ce50f1fef8c192e912e7c2b0d039209a7b676829abe13f4fc86b5ea20bc Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->
Extracting keywords, language, and embeddings
Use this for the enrichment surface around extraction: statistical keyword extraction, language detection, and vector embeddings. Keywords and language detection ride along with extraction and land on the result; embeddings are produced by a dedicated `embed` command.
Keywords (YAKE / RAKE)
Keyword extraction is configured via the `[keywords]` config block (or inline JSON) — there is no single `--keywords` CLI flag. When enabled, extracted keywords appear on `result.extracted_keywords` (`extractedKeywords` in Node.js; the CLI JSON field is `extracted_keywords`). Two algorithms are available:
- **YAKE** (`"yake"`) — statistical, unsupervised single-document
extraction. Good general default.
- **RAKE** (`"rake"`) — co-occurrence / phrase-based. Favors multi-word
key phrases.
> Feature-gated: keyword extraction requires the CLI to be built with the > `keywords-yake` and/or `keywords-rake` Cargo features (both are in the > default/`full` build). If the CLI was built without them, the `[keywords]` > config block is silently ignored — `result.extracted_keywords` simply stays empty > rather than erroring. The `"yake"` algorithm needs `keywords-yake`; `"rake"` > needs `keywords-rake`.
Enable via inline JSON on the CLI:
xberg extract paper.pdf --format json \
--config-json '{"keywords":{"algorithm":"yake","max_keywords":15,"language":"en"}}' \
| jq '.extracted_keywords'Or in a config file:
[keywords]
algorithm = "rake" # "yake" or "rake"
max_keywords = 10 # default 10
min_score = 0.0 # filter below this score (normalized 0.0-1.0 for both algorithms)
ngram_range = [1, 3] # unigrams..trigrams (default); config-file only
language = "en" # stopword language; omit to skip stopword filtering
xberg extract report.pdf --config xberg.toml --format json | jq '.extracted_keywords'
Field notes:
- `max_keywords` caps how many keywords are returned (default 10).
- `min_score` filters low-scoring keywords. Both YAKE and RAKE normalize
their scores to the `0.0`-`1.0` range with *higher-is-better*, so `min_score` retains keywords with `score >= min_score` identically for either algorithm.
- `ngram_range` is `[min, max]`: `[1,1]` unigrams only, `[1,2]` adds
bigrams, `[1,3]` (default) adds trigrams. Config-file only — it is not a field on the language bindings' `KeywordConfig`.
- `language` enables stopword filtering for that language; omit it to
disable stopword filtering entirely.
Language detection
Language detection is a real CLI flag: `--detect-language`. Detected languages appear on `result.detected_languages`:
xberg extract multilingual.pdf --detect-language true --format json \
| jq '.detected_languages'
In a config file it lives under `[language_detection]`:
[language_detection]
enabled = true
min_confidence = 0.8
detect_multiple = false
The CLI flag enables detection with `min_confidence = 0.8` and single-language mode; use the config block to detect multiple languages or tune confidence.
Embeddings (`embed` command)
The standalone `embed` command produces vector embeddings for text from `--text` (repeatable) or stdin. It does not run extraction — pipe extracted content in if you want document embeddings.
# Local ONNX preset model (default provider)
xberg embed --text "first passage" --text "second passage" --preset balanced
# Embed extracted document text
xberg extract report.pdf | xberg embed --preset quality
Presets for the local provider: `fast`, `balanced` (default), `quality`, `multilingual`. Output defaults to JSON (`--format json`).
`--provider` selects the embedding source:
| Provider | Flag | Notes | | -------- | ------------------------------------- | --------------------------------------------- | | `local` | `--preset <fast\|balanced\|quality\|multilingual>` | **Default.** ONNX model, no API key. | | `llm` | `--model <id>` `--api-key <key>` | liter-llm routing, e.g. `openai/text-embedding-3-small`. | | `plugin` | `--plugin <name>` | A backend pre-registered in-process via the plugin API. |
# Provider-hosted embeddings via an LLM
xberg embed --text "query text" \
--provider llm --model openai/text-embedding-3-small --api-key "$OPENAI_API_KEY"
Local embedding presets must be downloaded first if not cached. Pre-warm them with the cache command:
xberg cache warm --embedding-model balanced # one preset
xberg cache warm --all-embeddings # all available presets (currently 8)
Programmatic access
Keywords and detected languages live on the document in the result envelope:
from xberg import ExtractInput, extract, ExtractionConfig, KeywordConfig, KeywordAlgorithm
config = ExtractionConfig(
keywords=KeywordConfig(algorithm=KeywordAlgorithm.YAKE, max_keywords=15, language="en"),
)
result = await extract(ExtractInput(uri="paper.pdf"), config)
doc = result.results[0]
print(doc.extracted_keywords) # extracted keywords (when enabled)
print(doc.detected_languages) # detected languages (when enabled)See `references/python-api.md` and `references/configuration.md` in the sibling `xberg` skill for the keyword / language-detection config classes and the embedding presets.
##
Read more
name: extracting-keywords description: Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0da13ce50f1fef8c192e912e7c2b0d039209a7b676829abe13f4fc86b5ea20bc Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->
Extracting keywords, language, and embeddings
Use this for the enrichment surface around extraction: statistical keyword extraction, language detection, and vector embeddings. Keywords and language detection ride along with extraction and land on the result; embeddings are produced by a dedicated `embed` command.
Keywords (YAKE / RAKE)
Keyword extraction is configured via the `[keywords]` config block (or inline JSON) — there is no single `--keywords` CLI flag. When enabled, extracted keywords appear on `result.extracted_keywords` (`extractedKeywords` in Node.js; the CLI JSON field is `extracted_keywords`). Two algorithms are available:
- **YAKE** (`"yake"`) — statistical, unsupervised single-document
extraction. Good general default.
- **RAKE** (`"rake"`) — co-occurrence / phrase-based. Favors multi-word
key phrases.
> Feature-gated: keyword extraction requires the CLI to be built with the > `keywords-yake` and/or `keywords-rake` Cargo features (both are in the > default/`full` build). If the CLI was built without them, the `[keywords]` > config block is silently ignored — `result.extracted_keywords` simply stays empty > rather than erroring. The `"yake"` algorithm needs `keywords-yake`; `"rake"` > needs `keywords-rake`.
Enable via inline JSON on the CLI:
xberg extract paper.pdf --format json \
--config-json '{"keywords":{"algorithm":"yake","max_keywords":15,"language":"en"}}' \
| jq '.extracted_keywords'Or in a config file:
[keywords] algorithm = "rake" # "yake" or "rake" max_keywords = 10 # default 10 min_score = 0.0 # filter below this score (normalized 0.0-1.0 for both algorithms) ngram_range = [1, 3] # unigrams..trigrams (default); config-file only language = "en" # stopword language; omit to skip stopword filtering
xberg extract report.pdf --config xberg.toml --format json | jq '.extracted_keywords'
Field notes:
- `max_keywords` caps how many keywords are returned (default 10).
- `min_score` filters low-scoring keywords. Both YAKE and RAKE normalize
their scores to the `0.0`-`1.0` range with *higher-is-better*, so `min_score` retains keywords with `score >= min_score` identically for either algorithm.
- `ngram_range` is `[min, max]`: `[1,1]` unigrams only, `[1,2]` adds
bigrams, `[1,3]` (default) adds trigrams. Config-file only — it is not a field on the language bindings' `KeywordConfig`.
- `language` enables stopword filtering for that language; omit it to
disable stopword filtering entirely.
Language detection
Language detection is a real CLI flag: `--detect-language`. Detected languages appear on `result.detected_languages`:
xberg extract multilingual.pdf --detect-language true --format json \ | jq '.detected_languages'
In a config file it lives under `[language_detection]`:
[language_detection] enabled = true min_confidence = 0.8 detect_multiple = false
The CLI flag enables detection with `min_confidence = 0.8` and single-language mode; use the config block to detect multiple languages or tune confidence.
Embeddings (`embed` command)
The standalone `embed` command produces vector embeddings for text from `--text` (repeatable) or stdin. It does not run extraction — pipe extracted content in if you want document embeddings.
# Local ONNX preset model (default provider) xberg embed --text "first passage" --text "second passage" --preset balanced # Embed extracted document text xberg extract report.pdf | xberg embed --preset quality
Presets for the local provider: `fast`, `balanced` (default), `quality`, `multilingual`. Output defaults to JSON (`--format json`).
`--provider` selects the embedding source:
| Provider | Flag | Notes | | -------- | ------------------------------------- | --------------------------------------------- | | `local` | `--preset <fast\|balanced\|quality\|multilingual>` | **Default.** ONNX model, no API key. | | `llm` | `--model <id>` `--api-key <key>` | liter-llm routing, e.g. `openai/text-embedding-3-small`. | | `plugin` | `--plugin <name>` | A backend pre-registered in-process via the plugin API. |
# Provider-hosted embeddings via an LLM xberg embed --text "query text" \ --provider llm --model openai/text-embedding-3-small --api-key "$OPENAI_API_KEY"
Local embedding presets must be downloaded first if not cached. Pre-warm them with the cache command:
xberg cache warm --embedding-model balanced # one preset xberg cache warm --all-embeddings # all available presets (currently 8)
Programmatic access
Keywords and detected languages live on the document in the result envelope:
from xberg import ExtractInput, extract, ExtractionConfig, KeywordConfig, KeywordAlgorithm
config = ExtractionConfig(
keywords=KeywordConfig(algorithm=KeywordAlgorithm.YAKE, max_keywords=15, language="en"),
)
result = await extract(ExtractInput(uri="paper.pdf"), config)
doc = result.results[0]
print(doc.extracted_keywords) # extracted keywords (when enabled)
print(doc.detected_languages) # detected languages (when enabled)See `references/python-api.md` and `references/configuration.md` in the sibling `xberg` skill for the keyword / language-detection config classes and the embedding presets.
##
The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data.
Repo: xberg-io/xberg
Other skills on xberg.
- /batch-extraction
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.
Open skill - /chunking
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.
Open skill - /extracting-tables
Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown / JSON cells), and known limits.
Open skill - /extracting-with-ocr
Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
Open skill - /picking-a-format
Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
Open skill - /xberg
Extract text, tables, metadata, and images from 101 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg. Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async),
Open skill

