wigolo-agent
Autonomous data gathering across sources — plans search queries and URLs from a natural-language prompt, executes in parallel within a time budget, optionally…
Local-first multi-page crawl with sitemap, BFS, DFS, and URL-map strategies, anchor-fragment dedup, rate limiting, robots.txt respect, and automatic local cache population. Use when the user wants to index documentation, crawl a docs site, extract all pages under a path, or says
$ npx -y skills add KnockOutEZ/wigolo --skill wigolo-crawl --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/wigolo-crawlContext preview
The summary Claude sees to decide when to auto-load this skill.
Local-first multi-page crawl with sitemap, BFS, DFS, and URL-map strategies, anchor-fragment dedup, rate limiting, robots.txt respect, and automatic local cache population. Use when the user wants to index documentation, crawl a docs site, extract all pages under a path, or says
name: wigolo-crawl description: | Local-first multi-page crawl with sitemap, BFS, DFS, and URL-map strategies, anchor-fragment dedup, rate limiting, robots.txt respect, and automatic local cache population. Use when the user wants to index documentation, crawl a docs site, extract all pages under a path, or says "crawl", "index this site", "get all the docs", "bulk extract". Prefer when crawled pages should land in a reusable local cache for later `cache` / `find_similar` queries. license: AGPL-3.0-only metadata: author: KnockOutEZ version: 0.1.43-beta.2 homepage: https://github.com/KnockOutEZ/wigolo repository: https://github.com/KnockOutEZ/wigolo
Crawl sites with configurable strategy, depth, and rate limiting. All pages enter the local cache with embeddings.
// Crawl docs via sitemap (fastest, recommended for doc sites)
{ "url": "https://docs.example.com", "strategy": "sitemap", "max_pages": 30 }
// BFS crawl with scope filter
{ "url": "https://example.com", "strategy": "bfs", "max_depth": 3, "max_pages": 50, "include_patterns": ["^https://example\\.com/docs"] }
// URL discovery only (no content fetched — fastest for scoping)
{ "url": "https://example.com", "strategy": "map" }
// Authenticated crawl
{ "url": "https://app.example.com/docs", "strategy": "bfs", "use_auth": true, "max_pages": 20 }| Parameter | Type | Default | When to use | |-----------|------|---------|-------------| | `url` | string | required | Seed URL | | `strategy` | string | "bfs" | "sitemap" for doc sites, "map" for URL discovery only | | `max_depth` | number | 2 | How many link levels to follow | | `max_pages` | number | 20 | Hard cap on pages fetched | | `include_patterns` | string[] | none | Regex whitelist — ALWAYS add to stay in scope | | `exclude_patterns` | string[] | none | Regex blacklist | | `use_auth` | boolean | false | For authenticated sites | | `extract_links` | boolean | false | Return inter-page link graph | | `max_total_chars` | number | 100000 | Total char budget | | `max_tokens_out` | number | none | Token-budget cap (cl100k-base) | | `include_full_markdown` | boolean | false | Pages return evidence-only by default | | `citation_format` | string | "numbered" | "numbered" / "json" / "anthropic_tags" |
Pages are dedup'd by canonical URL (anchor-fragment aware: `/intro` and `/intro#install` collapse to one entry).
All crawled pages enter the local cache with embeddings. This means:
**Crawl first, then use `cache` and `find_similar` for subsequent lookups.**
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
Repo: KnockOutEZ/wigolo
Autonomous data gathering across sources — plans search queries and URLs from a natural-language prompt, executes in parallel within a time budget, optionally…
Local-first knowledge cache — full-text and hybrid semantic search over every page wigolo has already fetched, crawled, or searched. Use before any web…
Compare two versions of a page and see exactly what changed — a live URL against its cached copy, two URLs, or two markdown blobs. Section-level hunks, word-…
Local-first structured extraction from any webpage — tables, definition lists, key-value pairs, JSON-LD, microdata, chart hints (SVG titles / aria-labels /…
Local-first URL fetch with clean markdown, structured metadata, JS-rendered SPA support, authenticated browser sessions, PDFs, and content change detection.…
Hybrid semantic discovery — fuses embeddings + keyword search + live web search via 3-way Reciprocal Rank Fusion. Use when the user has a good source and wants…