/doc-scraper
Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills.
$ npx -y skills add jmagly/aiwg --skill doc-scraper --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/doc-scraper
Context preview
The summary Claude sees to decide when to auto-load this skill.
Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills.
SKILL.md
doc-scraper.SKILL.mdnamespace: aiwg
name: doc-scraper
description: Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills.
tools: Read, Write, Bash, WebFetch
platforms: [all]
Documentation Scraper Skill
Purpose
Single responsibility: Convert documentation websites into organized, categorized reference files suitable for Claude skills or offline archives. (BP-4)
Grounding Checkpoint (Archetype 1 Mitigation)
Before executing, VERIFY:
- [ ] Target URL is accessible (test with `curl -I`)
- [ ] Documentation structure is identifiable (inspect page for content selectors)
- [ ] Output directory is writable
- [ ] Rate limiting requirements are known (check robots.txt)
**DO NOT proceed without verification. Inspect before scraping.**
Uncertainty Escalation (Archetype 2 Mitigation)
ASK USER instead of guessing when:
- Content selector is ambiguous (multiple `<article>` or `<main>` elements)
- URL patterns unclear (can't determine include/exclude rules)
- Category mapping uncertain (content doesn't fit predefined categories)
- Rate limiting unknown (no robots.txt, unclear ToS)
**NEVER substitute missing configuration with assumptions.**
Context Scope (Archetype 3 Mitigation)
| Context Type | Included | Excluded | |--------------|----------|----------| | RELEVANT | Target URL, selectors, output path | Unrelated documentation | | PERIPHERAL | Similar site examples for selector hints | Historical scrape data | | DISTRACTOR | Other projects, unrelated URLs | Previous failed attempts |
Workflow Steps
Step 1: Verify Target (Grounding)
# Test URL accessibility
curl -I <target-url>
# Check robots.txt
curl <base-url>/robots.txt
# Inspect page structure (use browser dev tools or fetch sample)
Step 2: Create Configuration
Generate scraper config based on inspection:
{
"name": "skill-name",
"description": "When to use this skill",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article",
"title": "h1",
"code_blocks": "pre code"
},
"url_patterns": {
"include": ["/docs", "/guide", "/api"],
"exclude": ["/blog", "/changelog", "/releases"]
},
"categories": {
"getting_started": ["intro", "quickstart", "installation"],
"api_reference": ["api", "reference", "methods"],
"guides": ["guide", "tutorial", "how-to"]
},
"rate_limit": 0.5,
"max_pages": 500
}Step 3: Execute Scraping
**Option A: With skill-seekers (if installed)**
# Verify skill-seekers is available
pip show skill-seekers
# Run scraper
skill-seekers scrape --config config.json
# For large docs, use async mode
skill-seekers scrape --config config.json --async --workers 8
**Option B: Manual scraping guidance**
1. Use sitemap.xml or crawl starting URL 2. Extract content using configured selectors 3. Categorize pages based on URL patterns and keywords 4. Save to organized directory structure
Step 4: Validate Output
# Check output structure
ls -la output/<skill-name>/
# Verify content quality
head -50 output/<skill-name>/references/index.md
# Count extracted pages
find output/<skill-name>_data/pages -name "*.json" | wc -l
Recovery Protocol (Archetype 4 Mitigation)
On error:
1. **PAUSE** - Stop scraping, preserve already-fetched pages 2. **DIAGNOSE** - Check error type:
- `Connection error` → Verify URL, check network
- `Selector not found` → Re-inspect page structure
- `Rate limited` → Increase delay, reduce workers
- `Memory/disk` → Reduce batch size, clear temp files
3. **ADAPT** - Adjust configuration based on diagnosis 4. **RETRY** - Resume from checkpoint (max 3 attempts) 5. **ESCALATE** - Ask user for guidance
Checkpoint Support
State saved to: `.aiwg/working/checkpoints/doc-scraper/`
Resume interrupted scrape:
skill-seekers scrape --config config.json --resume
Clear checkpoint and start fresh:
skill-seekers scrape --config config.json --fresh
Output Structure
output/<skill-name>/
├── SKILL.md # Main skill description
├── references/ # Categorized documentation
│ ├── index.md # Category index
│ ├── getting_started.md
│ ├── api_reference.md
│ └── guides.md
├── scripts/ # (empty, for user additions)
└── assets/ # (empty, for user additions)
output/<skill-name>_data/
├── pages/ # Raw scraped JSON (one per page)
└── summary.json # Scrape statistics
Configuration Templates
Minimal Config
{
"name": "myframework",
"base_url": "https://docs.example.com/",
"max_pages": 100
}Full Config
{
"name": "myframework",
"description": "MyFramework documentation for building web apps",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article, main, div[role='main']",
"title": "h1, .title",
"code_blocks": "pre code, .highlight code",
"navigation": "nav, .sidebar"
},
"url_patterns": {
"include": ["/docs/", "/api/", "/guide/"],
"exclude": ["/blog/", "/changelog/", "/v1/", "/v2/"]
},
"categories": {
"getting_started": ["intro", "quickstart", "install", "setup"],
"concepts": ["concept", "overview", "architecture"],
"api": ["api", "reference", "method", "function"],
"guides": ["guide", "tutorial", "how-to", "example"],
"advanced": ["advanced", "internals", "customize"]
},
"rate_limit": 0.5,
"max_pages": 1000,
"checkpoint": {
"enabled": true,
"interval": 100
}
}Troubleshooting
| Issue | Diagnosis | Solution | |-------|-----------|----------| | No content extracted | Selector mismatch | Inspect page, update `main_content` selector | | Wrong pages scraped | URL pattern issue | Check `include`/`exclude` patterns | | Rate limited | Too aggressive | Increase `rate_limit` to 1.0+ seconds | | Memory issues | Too many pages |
Read more
namespace: aiwg name: doc-scraper description: Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills. tools: Read, Write, Bash, WebFetch platforms: [all]
Documentation Scraper Skill
Purpose
Single responsibility: Convert documentation websites into organized, categorized reference files suitable for Claude skills or offline archives. (BP-4)
Grounding Checkpoint (Archetype 1 Mitigation)
Before executing, VERIFY:
- [ ] Target URL is accessible (test with `curl -I`)
- [ ] Documentation structure is identifiable (inspect page for content selectors)
- [ ] Output directory is writable
- [ ] Rate limiting requirements are known (check robots.txt)
**DO NOT proceed without verification. Inspect before scraping.**
Uncertainty Escalation (Archetype 2 Mitigation)
ASK USER instead of guessing when:
- Content selector is ambiguous (multiple `<article>` or `<main>` elements)
- URL patterns unclear (can't determine include/exclude rules)
- Category mapping uncertain (content doesn't fit predefined categories)
- Rate limiting unknown (no robots.txt, unclear ToS)
**NEVER substitute missing configuration with assumptions.**
Context Scope (Archetype 3 Mitigation)
| Context Type | Included | Excluded | |--------------|----------|----------| | RELEVANT | Target URL, selectors, output path | Unrelated documentation | | PERIPHERAL | Similar site examples for selector hints | Historical scrape data | | DISTRACTOR | Other projects, unrelated URLs | Previous failed attempts |
Workflow Steps
Step 1: Verify Target (Grounding)
# Test URL accessibility curl -I <target-url> # Check robots.txt curl <base-url>/robots.txt # Inspect page structure (use browser dev tools or fetch sample)
Step 2: Create Configuration
Generate scraper config based on inspection:
{
"name": "skill-name",
"description": "When to use this skill",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article",
"title": "h1",
"code_blocks": "pre code"
},
"url_patterns": {
"include": ["/docs", "/guide", "/api"],
"exclude": ["/blog", "/changelog", "/releases"]
},
"categories": {
"getting_started": ["intro", "quickstart", "installation"],
"api_reference": ["api", "reference", "methods"],
"guides": ["guide", "tutorial", "how-to"]
},
"rate_limit": 0.5,
"max_pages": 500
}Step 3: Execute Scraping
**Option A: With skill-seekers (if installed)**
# Verify skill-seekers is available pip show skill-seekers # Run scraper skill-seekers scrape --config config.json # For large docs, use async mode skill-seekers scrape --config config.json --async --workers 8
**Option B: Manual scraping guidance**
1. Use sitemap.xml or crawl starting URL 2. Extract content using configured selectors 3. Categorize pages based on URL patterns and keywords 4. Save to organized directory structure
Step 4: Validate Output
# Check output structure ls -la output/<skill-name>/ # Verify content quality head -50 output/<skill-name>/references/index.md # Count extracted pages find output/<skill-name>_data/pages -name "*.json" | wc -l
Recovery Protocol (Archetype 4 Mitigation)
On error:
1. **PAUSE** - Stop scraping, preserve already-fetched pages 2. **DIAGNOSE** - Check error type:
- `Connection error` → Verify URL, check network
- `Selector not found` → Re-inspect page structure
- `Rate limited` → Increase delay, reduce workers
- `Memory/disk` → Reduce batch size, clear temp files
3. **ADAPT** - Adjust configuration based on diagnosis 4. **RETRY** - Resume from checkpoint (max 3 attempts) 5. **ESCALATE** - Ask user for guidance
Checkpoint Support
State saved to: `.aiwg/working/checkpoints/doc-scraper/`
Resume interrupted scrape:
skill-seekers scrape --config config.json --resume
Clear checkpoint and start fresh:
skill-seekers scrape --config config.json --fresh
Output Structure
output/<skill-name>/ ├── SKILL.md # Main skill description ├── references/ # Categorized documentation │ ├── index.md # Category index │ ├── getting_started.md │ ├── api_reference.md │ └── guides.md ├── scripts/ # (empty, for user additions) └── assets/ # (empty, for user additions) output/<skill-name>_data/ ├── pages/ # Raw scraped JSON (one per page) └── summary.json # Scrape statistics
Configuration Templates
Minimal Config
{
"name": "myframework",
"base_url": "https://docs.example.com/",
"max_pages": 100
}Full Config
{
"name": "myframework",
"description": "MyFramework documentation for building web apps",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article, main, div[role='main']",
"title": "h1, .title",
"code_blocks": "pre code, .highlight code",
"navigation": "nav, .sidebar"
},
"url_patterns": {
"include": ["/docs/", "/api/", "/guide/"],
"exclude": ["/blog/", "/changelog/", "/v1/", "/v2/"]
},
"categories": {
"getting_started": ["intro", "quickstart", "install", "setup"],
"concepts": ["concept", "overview", "architecture"],
"api": ["api", "reference", "method", "function"],
"guides": ["guide", "tutorial", "how-to", "example"],
"advanced": ["advanced", "internals", "customize"]
},
"rate_limit": 0.5,
"max_pages": 1000,
"checkpoint": {
"enabled": true,
"interval": 100
}
}Troubleshooting
| Issue | Diagnosis | Solution | |-------|-----------|----------| | No content extracted | Selector mismatch | Inspect page, update `main_content` selector | | Wrong pages scraped | URL pattern issue | Check `include`/`exclude` patterns | | Rate limited | Too aggressive | Increase `rate_limit` to 1.0+ seconds | | Memory issues | Too many pages |
Multi-agent AI framework for Claude Code, Copilot, Cursor, Warp, and 6 more platforms 200+ agents, 109+ CLI commands, 400+ deployable agent/skill/command/rule artifacts, 8 core frameworks, 32 addons, and a 40-plugin Claude Code marketplace.
Repo: jmagly/aiwg
Other skills on aiwg.
- /agent-loop-ext
Crash-resilient external agent loop with state persistence and CI/CD integration
Open skill - /agent-loop
Detect requests for iterative autonomous agent loops and route to the appropriate loop executor
Open skill - /auto-test-execution
Automatically execute tests when code-generating agents modify source files, enforcing the execute-before-return pattern
Open skill - /cross-task-learner
Enable agent loops to learn from similar past tasks and share patterns across loops
Open skill - /debug-memory
Query and manage the executable feedback debug memory
Open skill - /execute-feedback
Execute tests on generated code and iterate until passing
Open skill

