Skip to content
Research
Command

/ingest-collection

Bulk-ingest source collections such as Git doc repos, MediaWiki sources, CSV/JSON message archives, and Wayback CDX snapshots into raw sources.

BOOST
From plugin
nvk-llm-wiki
1.4k28 skills28 commands
Install
> /plugin marketplace add nvk/llm-wiki
> /plugin install wiki@llm-wiki

How it fires

How this command gets triggered: by you, by Claude, or both.

  • Fires itselfClaude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/ingest-collection

Context preview

What this command does when you run it.

Bulk-ingest source collections such as Git doc repos, MediaWiki sources, CSV/JSON message archives, and Wayback CDX snapshots into raw sources.

Command definition

ingest-collection.md
description: "Bulk-ingest source collections such as Git doc repos, MediaWiki sources, CSV/JSON message archives, and Wayback CDX snapshots into raw sources."
argument-hint: "<repo-url|repo-path|mediawiki-url|dump.xml[.bz2|.gz]|csv-json-path|cdx-url|archived-url> [--adapter auto|git|mediawiki-dump|mediawiki-api|csv-messages|wayback-cdx] [--wiki <name>] [--local] [--new-topic <name>] [--limit <N>] [--namespace <id>] [--include <pattern>] [--exclude <pattern>] [--from YYYYMMDD] [--to YYYYMMDD] [--dry-run] [--compile] [--include-archived]"
allowed-tools: Read, Write, Edit, Glob, Grep, Bash(ls:*), Bash(wc:*), Bash(date:*), Bash(mkdir:*), Bash(mv:*), Bash(cp:*), Bash(rm:*), Bash(basename:*), Bash(find:*), Bash(git:*), Bash(curl:*), Bash(python3:*), Bash(bunzip2:*), Bash(gunzip:*)

Your task

Bulk-ingest a collection into the wiki as immutable raw sources. A collection is a bounded upstream corpus, not a single article: a Git repository full of specs, a BIP repository, a MediaWiki XML dump/API site, a CSV/JSON message archive, or a Wayback CDX snapshot set. Do **not** compile one wiki article per upstream page by default. First preserve raw sources with provenance, then compile synthesized topic/concept/reference articles.

**Resolve the wiki.** Follow the same resolution flow as `/wiki:ingest`: 1. Read `$HOME/.config/llm-wiki/config.json`. If it has `hub_path`, expand leading `~` only and prefer that path; use `resolved_path` only as a fallback cache when the expanded `hub_path` is unavailable and `resolved_path` is initialized. If config has only `resolved_path`, use it. If the configured path can be statted but reading `wikis.json` or listing `topics/` fails with `Operation not permitted`, stop and ask the user to grant Full Disk Access/iCloud Drive access to the launcher; do not fall back to `~/wiki` or `resolved_path`. Do not write machine-specific `resolved_path` into shared configs. 2. If no config → read `$HOME/wiki/_index.md`. If it exists → HUB = `$HOME/wiki`. If nothing found, ask the user where to create the wiki. 3. Wiki location, first match: `--local` → `.wiki/`; `--wiki <name>` → `HUB/wikis.json` lookup with portable path resolution (`<HUB>`, `~`, absolute, or HUB-relative); if the registry path is stale, fall back to `HUB/topics/<name>`; current directory has `.wiki/` → use it; else → HUB. 4. If `<wiki>/_index.md` is missing and `--new-topic <name>` is set, create the topic wiki using the init protocol before ingesting. If no wiki exists and no `--new-topic`, stop and ask for a target wiki.

Archive rule: collection ingest skips archived topic wikis by default. If `--wiki <name>` resolves to `status: archived` or a path under `topics/.archive/`, stop and ask the user to restore it with `/wiki:archive restore <name>` or rerun with `--include-archived`. When explicitly included, write only inside that archived topic path and keep it archived. Auto-classification and `--new-topic` collision checks should treat archived topics as unavailable unless the user explicitly restores or includes archived context.

Read `skills/wiki-manager/references/ingestion.md` and `skills/wiki-manager/references/wiki-structure.md`, then follow the collection ingestion protocol.

Inventory and dataset awareness

Before writing a large collection, be explicit about fit:

  • If the user only wants to remember a possible collection for later, create or

suggest one inventory `corpus` record instead of ingesting.

  • If the collection is row-like data that should be queried in place, create or

suggest a dataset manifest plus one linked inventory record.

  • If the user is about to ingest hundreds of child sources, show the collection

manifest shape, estimated child count, and any inventory/dataset companion record before asking for confirmation.

After ingesting a collection, if a matching inventory record exists, link the raw collection manifest from that record and report the recommended status/next action update.

Parse arguments

  • **Source**: repo URL/path, MediaWiki site URL, dump file/URL, CSV/TSV/JSON/JSONL path/URL, CDX API URL, or original URL to query through the Wayback CDX API.
  • **--adapter**: `auto` default. Supported: `git`, `mediawiki-dump`, `mediawiki-api`, `csv-messages`, `wayback-cdx`.
  • **--limit <N>**: maximum child sources to ingest. Useful for API imports and dry runs.
  • **--namespace <id>**: MediaWiki namespace. Default `0`.
  • **--include <pattern> / --exclude <pattern>**: filter upstream paths, page titles, message fields, or original snapshot URLs. Treat as shell globs for Git paths and regex for MediaWiki titles, message text, and Wayback original URLs unless the user specifies otherwise.
  • **--from YYYYMMDD / --to YYYYMMDD**: bound Wayback CDX capture timestamps.
  • **--dry-run**: list what would be ingested, write nothing.
  • **--compile**: after raw ingestion, run the normal compile workflow with collection-aware clustering.
  • **--include-archived**: explicitly allow ingestion into an archived target

wiki. Keep the topic archived.

Adapter detection

Use `--adapter` if supplied. Otherwise:

1. Source ending in `.xml`, `.xml.bz2`, or `.xml.gz` → `mediawiki-dump`. 2. Source contains `github.com/`, `gitlab.com/`, ends in `.git`, or is a local directory with `.git/` → `git`. 3. Source ending in `.csv`, `.tsv`, `.json`, or `.jsonl` with message-like fields, or user asks for per-message markdown → `csv-messages`. 4. Source contains `web.archive.org/cdx`, user says "Wayback"/"CDX", or source is an original URL plus snapshot/archive intent → `wayback-cdx`. 5. Source URL contains `/wiki/`, `/w/`, or has a reachable `api.php` endpoint → `mediawiki-api`. 6. If still ambiguous, ask the user to choose `git`, `mediawiki-dump`, `mediawiki-api`, `csv-messages`, or `wayback-cdx`.

Never recursively crawl HTML pages as a collection import. Use structured upstream APIs, repository files, official dumps, dataset rows, or CDX inventories.

Shared collection flow

1. Deri

Read more
Ships withnvk-llm-wiki

LLM-compiled knowledge bases for any AI agent. Parallel multi-agent research, thesis-driven investigation, source ingestion, wiki compilation, querying, and artifact generation.

Get the whole plugin
Stats
1,394
Stars
131
Forks
Active
Maintenance
Python
Language
MIT
License
6d ago
Last commit
6mo ago
Created
1d ago
Added

Repo: nvk/llm-wiki

Other commands on nvk-llm-wiki.