agent-onboarding
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content). Use when the user wants "live research", to "research <topic> deeply", "research the latest on", "write a report
$ npx -y skills add brightdata/skills --skill live-research --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/live-researchContext preview
The summary Claude sees to decide when to auto-load this skill.
Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content). Use when the user wants "live research", to "research <topic> deeply", "research the latest on", "write a report
name: live-research description: | Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content). Use when the user wants "live research", to "research <topic> deeply", "research the latest on", "write a report on", "give me a briefing / literature review / market scan", "find and synthesize everything about", or otherwise wants a synthesized, source-grounded answer rather than a list of links. Decomposes the question into multiple intent-ranked Discover queries, pulls page content, deduplicates and ranks by relevance, then synthesizes a structured brief with inline citations. Built on the `discover-api` skill. For competitor-specific intel use `competitive-intel`; for social/brand sentiment use `brand-listening`; for a retrieval *system* (not a one-off report) use `rag-pipeline`. metadata: author: Bright Data version: "1.0"
Turn one research question into a **cited, synthesized brief** by fanning out intent-ranked Discover queries, reading the best sources, and writing up findings with inline citations. This is a *workflow* on top of the **`discover-api`** skill — read that for the API mechanics, modes, and parameters.
Use this when the deliverable is *understanding* (a report/briefing), not a link list (that's `search`/`discover-api`) and not a standing system (that's `rag-pipeline`).
Discover must be reachable. Quick check (CLI path):
command -v bdata >/dev/null 2>&1 || echo "CLI missing — see bright-data-best-practices/references/cli-setup.md" bdata zones >/dev/null 2>&1 || echo "not authenticated — run: bdata login"
(SDK/REST paths just need `BRIGHTDATA_API_TOKEN`.)
If the question is broad or ambiguous, ask 2–3 clarifying questions before spending API calls: time horizon, geography/market, depth, and what decision the research supports. A sharp scope is what makes the `intent` parameters good.
Break the topic into 4–8 angles (definitions, key players, mechanisms, evidence, counter-evidence, recent developments, risks). Each angle becomes one Discover call with its own tailored `intent`. This beats one broad query — `num_results` is capped at 20, so coverage comes from *breadth of queries*, not one big call.
# one call per angle; --include-content so you read sources in the same pass bdata discover "stablecoin regulation 2026" \ --intent "recent regulatory actions and proposed legislation, primary sources" \ --include-content --num-results 15 -o angle_regulation.json & bdata discover "stablecoin reserve transparency" \ --intent "audits, attestations, reserve composition disclosures" \ --include-content --num-results 15 -o angle_reserves.json & wait
For *maximum* coverage on a hard topic, use the raw REST flow with `"mode":"deep"` (see `discover-api`) — `deep` is exhaustive but slower and REST-only.
# VERIFIED: this is the correct merge. `jq -s 'add | unique_by(.link)'` does NOT work —
# each file is {results:[...]}, so you must flatten .results[] first.
jq -s '
[ .[].results[] ] # flatten results from all files
| unique_by(.link) # dedup by URL
| map(select(
.content != null
and (.content | length) > 200 # drop empty / 404 stubs
and ((.content | test("just a moment|captcha|access denied|cf-browser|page not found|post not found"; "i")) | not)
))
| sort_by(-.relevance_score)
' angle_*.json > corpus.json
echo "kept $(jq length corpus.json) sources"**Or just run the helper** (same logic, tested): `scripts/merge_corpus.sh -o corpus.json angle_*.json` (`-m <n>` sets the min content length). Copying the jq by hand is error-prone — prefer the script.
> Note: with `--include-content`, the leading part of `content` is usually page > nav/boilerplate (menus, logos). When extracting claims (Step 5), skip past the > chrome to the article body.
From each kept source's `content`, pull the specific claims, numbers, dates, and quotes that answer a sub-question. Track which URL each claim came from — you'll cite it.
Write the structured brief (template in `references/brief-template.md`). Every non-obvious claim gets an inline citation `[n]` mapping to a numbered source list. Note disagreements between sources rather than averaging them away.
Repo: brightdata/skills
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Social listening and brand reputation research using Bright Data's web scraping…
Debug Bright Data Scraping Browser sessions using the Browser Sessions API. Use this skill…
Build production-ready Bright Data integrations with best practices baked in. Reference…
Bright Data MCP handles ALL web data operations. Replaces WebFetch, WebSearch, and all…
Guide for using the Bright Data CLI (`brightdata` / `bdata`) to scrape websites, search the…