agent-onboarding
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (`bdata scrape`). Use when the user wants to fetch a page, extract content from a list of URLs, or crawl paginated listings. Hands off to `data-feeds` for supported platforms (Amazon, LinkedIn, TikTok,
$ npx -y skills add brightdata/skills --skill scrape --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/scrapeContext preview
The summary Claude sees to decide when to auto-load this skill.
Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (`bdata scrape`). Use when the user wants to fetch a page, extract content from a list of URLs, or crawl paginated listings. Hands off to `data-feeds` for supported platforms (Amazon, LinkedIn, TikTok,
name: scrape description: Scrape web content as clean markdown/HTML/JSON via the Bright Data CLI (`bdata scrape`). Use when the user wants to fetch a page, extract content from a list of URLs, or crawl paginated listings. Hands off to `data-feeds` for supported platforms (Amazon, LinkedIn, TikTok, Instagram, YouTube, Reddit, etc.) and to `search` when URLs must be discovered first. Requires the Bright Data CLI; proactively guides install + login if missing.
Get clean content (markdown, HTML, JSON, screenshot) from one or more URLs via the Bright Data CLI. This skill owns the "fetch raw or lightly-structured content" job. For platform-specific structured data (Amazon, LinkedIn, TikTok, etc.), **stop and use `data-feeds` instead** — you'll get clean JSON without selector logic.
Before any scrape, verify the CLI is installed and authenticated:
if ! command -v bdata >/dev/null 2>&1; then
echo "bdata CLI not installed — see bright-data-best-practices/references/cli-setup.md"
elif ! bdata zones >/dev/null 2>&1; then
echo "bdata not authenticated — run: bdata login (or: bdata login --device for SSH)"
fiIf either check fails, halt and route the user to `skills/bright-data-best-practices/references/cli-setup.md`. Do not attempt the legacy `curl` fallback silently — ask the user first.
| Situation | Action | |---|---| | Single URL | `bdata scrape <url> -f markdown` | | Small list (≤ ~20 URLs) | shell loop, 1 at a time (see `references/patterns.md`) | | Larger list (dozens+) | `xargs -P 4` with parallelism cap (see `references/patterns.md`) | | Paginated listing | scrape page 1 → extract next-page URL → append → repeat (see `references/examples.md`) | | JS-heavy / login-gated / interaction-required | escalate to `bdata browser` (see `brightdata-cli` skill) | | Amazon, LinkedIn, TikTok, Instagram, YouTube, Reddit, … | **stop — hand off to `data-feeds`** | | No URL yet, just a topic | **hand off to `search`** |
Core commands:
# Clean markdown (default) bdata scrape "https://example.com/article" -f markdown -o article.md # Raw HTML (when you need the DOM) bdata scrape "https://example.com" -f html -o page.html # Structured JSON (when the Unlocker returns parsed fields) bdata scrape "https://example.com" -f json --pretty -o page.json # Visual snapshot (saves PNG) bdata scrape "https://example.com" -f screenshot -o page.png # Geo-targeted (override the exit country) bdata scrape "https://example.com" --country de -f markdown
Full flag reference: [`references/flags.md`](references/flags.md).
1. **Non-empty output:** `test -s "$out_path"` — or, for stdout, at least 200 bytes of content. 2. **Not a block page** — grep the output for any of these signatures (case-insensitive):
3. **Expected markers present** for the task: e.g., a product page should contain a price pattern (`\$\d`); an article should contain at least one `<h1>` or `# ` heading. 4. **On failure, escalation ladder:**
Do not report success until all checks above pass.
Repo: brightdata/skills
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Social listening and brand reputation research using Bright Data's web scraping…
Debug Bright Data Scraping Browser sessions using the Browser Sessions API. Use this skill…
Build production-ready Bright Data integrations with best practices baked in. Reference…
Bright Data MCP handles ALL web data operations. Replaces WebFetch, WebSearch, and all…
Guide for using the Bright Data CLI (`brightdata` / `bdata`) to scrape websites, search the…