agent-loop-ext
Crash-resilient external agent loop with state persistence and CI/CD integration
Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills.
$ npx -y skills add jmagly/aiwg --skill doc-scraper --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/doc-scraperContext preview
The summary Claude sees to decide when to auto-load this skill.
Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills.
namespace: aiwg name: doc-scraper description: Scrape documentation websites into organized reference files. Use when converting docs sites to searchable references or building Claude skills. tools: Read, Write, Bash, WebFetch platforms: [all]
Single responsibility: Convert documentation websites into organized, categorized reference files suitable for Claude skills or offline archives. (BP-4)
Before executing, VERIFY:
**DO NOT proceed without verification. Inspect before scraping.**
ASK USER instead of guessing when:
**NEVER substitute missing configuration with assumptions.**
| Context Type | Included | Excluded | |--------------|----------|----------| | RELEVANT | Target URL, selectors, output path | Unrelated documentation | | PERIPHERAL | Similar site examples for selector hints | Historical scrape data | | DISTRACTOR | Other projects, unrelated URLs | Previous failed attempts |
# Test URL accessibility curl -I <target-url> # Check robots.txt curl <base-url>/robots.txt # Inspect page structure (use browser dev tools or fetch sample)
Generate scraper config based on inspection:
{
"name": "skill-name",
"description": "When to use this skill",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article",
"title": "h1",
"code_blocks": "pre code"
},
"url_patterns": {
"include": ["/docs", "/guide", "/api"],
"exclude": ["/blog", "/changelog", "/releases"]
},
"categories": {
"getting_started": ["intro", "quickstart", "installation"],
"api_reference": ["api", "reference", "methods"],
"guides": ["guide", "tutorial", "how-to"]
},
"rate_limit": 0.5,
"max_pages": 500
}**Option A: With skill-seekers (if installed)**
# Verify skill-seekers is available pip show skill-seekers # Run scraper skill-seekers scrape --config config.json # For large docs, use async mode skill-seekers scrape --config config.json --async --workers 8
**Option B: Manual scraping guidance**
1. Use sitemap.xml or crawl starting URL 2. Extract content using configured selectors 3. Categorize pages based on URL patterns and keywords 4. Save to organized directory structure
# Check output structure ls -la output/<skill-name>/ # Verify content quality head -50 output/<skill-name>/references/index.md # Count extracted pages find output/<skill-name>_data/pages -name "*.json" | wc -l
On error:
1. **PAUSE** - Stop scraping, preserve already-fetched pages 2. **DIAGNOSE** - Check error type:
3. **ADAPT** - Adjust configuration based on diagnosis 4. **RETRY** - Resume from checkpoint (max 3 attempts) 5. **ESCALATE** - Ask user for guidance
State saved to: `.aiwg/working/checkpoints/doc-scraper/`
Resume interrupted scrape:
skill-seekers scrape --config config.json --resume
Clear checkpoint and start fresh:
skill-seekers scrape --config config.json --fresh
output/<skill-name>/ ├── SKILL.md # Main skill description ├── references/ # Categorized documentation │ ├── index.md # Category index │ ├── getting_started.md │ ├── api_reference.md │ └── guides.md ├── scripts/ # (empty, for user additions) └── assets/ # (empty, for user additions) output/<skill-name>_data/ ├── pages/ # Raw scraped JSON (one per page) └── summary.json # Scrape statistics
{
"name": "myframework",
"base_url": "https://docs.example.com/",
"max_pages": 100
}{
"name": "myframework",
"description": "MyFramework documentation for building web apps",
"base_url": "https://docs.example.com/",
"selectors": {
"main_content": "article, main, div[role='main']",
"title": "h1, .title",
"code_blocks": "pre code, .highlight code",
"navigation": "nav, .sidebar"
},
"url_patterns": {
"include": ["/docs/", "/api/", "/guide/"],
"exclude": ["/blog/", "/changelog/", "/v1/", "/v2/"]
},
"categories": {
"getting_started": ["intro", "quickstart", "install", "setup"],
"concepts": ["concept", "overview", "architecture"],
"api": ["api", "reference", "method", "function"],
"guides": ["guide", "tutorial", "how-to", "example"],
"advanced": ["advanced", "internals", "customize"]
},
"rate_limit": 0.5,
"max_pages": 1000,
"checkpoint": {
"enabled": true,
"interval": 100
}
}| Issue | Diagnosis | Solution | |-------|-----------|----------| | No content extracted | Selector mismatch | Inspect page, update `main_content` selector | | Wrong pages scraped | URL pattern issue | Check `include`/`exclude` patterns | | Rate limited | Too aggressive | Increase `rate_limit` to 1.0+ seconds | | Memory issues | Too many pages |
Reusable project context and specialist workflows for the AI tools you already use. Plan software, coordinate specialist reviews, prepare campaigns, investigate incidents, organize research, curate media, and maintain operational knowledge.
Repo: jmagly/aiwg
Crash-resilient external agent loop with state persistence and CI/CD integration
Detect requests for iterative autonomous agent loops and route to the appropriate loop executor
Automatically execute tests when code-generating agents modify source files, enforcing the execute-before-return pattern
Enable agent loops to learn from similar past tasks and share patterns across loops
Query and manage the executable feedback debug memory
Execute tests on generated code and iterate until passing