agent-onboarding
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Trigger a discovery job and retrieve ranked results (link, title, description, relevance_score) with optional parsed page content. Use when the user wants
$ npx -y skills add brightdata/skills --skill discover-api --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/discover-apiContext preview
The summary Claude sees to decide when to auto-load this skill.
Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Trigger a discovery job and retrieve ranked results (link, title, description, relevance_score) with optional parsed page content. Use when the user wants
name: discover-api description: | Use Bright Data's Discover API — intent-ranked, AI-relevance-scored web search at scale (not keyword SERP). Trigger a discovery job and retrieve ranked results (link, title, description, relevance_score) with optional parsed page content. Use when the user wants semantic/intent-based web search, "find pages about <topic> that match <goal>", web-grounded retrieval for an LLM, or results filtered by relevance rather than raw keyword rank. Covers the REST API (POST/GET /discover), the CLI (`bdata discover`), and the Python/JS SDKs (`client.discover`), including the standard/zeroRanking/deep/fast modes. This is the foundation skill for `live-research` and `rag-pipeline`. For keyword SERP use `search`; for structured platform data use `data-feeds`. metadata: author: Bright Data version: "1.0" documentation: https://docs.brightdata.com/api-reference/discover/overview
Discover is **intent-ranked semantic web search**. You give it a `query` plus an `intent`, and it returns results scored by AI relevance — optionally with the full parsed page content. It is the right primitive when result *quality/relevance* matters more than raw keyword rank, and the building block for retrieval (RAG), research, and knowledge-base pipelines.
**Discover vs. the neighbors:**
**`live-research`** or **`rag-pipeline`** (both call this API).
1. **Trigger** a job → you get a `task_id`. 2. **Poll** with the `task_id` until `status` is `"done"` (intermediate: `"processing"`). 3. Read `results[]`.
The CLI and SDKs do the trigger+poll for you; the raw REST flow is shown below for when you need parameters the wrappers don't expose (notably `mode`).
| You are… | Use | |---|---| | In a terminal, one-off or scripted | CLI: `bdata discover` | | Writing Node/TS code | JS SDK: `client.discover()` — see `js-sdk-best-practices` | | Writing Python code | Python SDK: `client.discover()` — see `python-sdk-best-practices` | | Need `mode` (deep/fast/zeroRanking) or `include_images` | **Raw REST** (wrappers don't expose these yet) |
Setup gate first:
command -v bdata >/dev/null 2>&1 || echo "CLI missing — see bright-data-best-practices/references/cli-setup.md" bdata zones >/dev/null 2>&1 || echo "not authenticated — run: bdata login"
# Intent-ranked discovery, JSON bdata discover "enterprise LLM platforms" \ --intent "vendor pages with pricing" \ --num-results 15 --json --pretty # With parsed page content in one pass (for RAG / research) bdata discover "webhook retry best practices" \ --include-content --num-results 10 -o results.json # Date-bounded bdata discover "react server components" \ --start-date 2025-01-01 --end-date 2025-12-31 --num-results 20 --json
Results live at `.results[]`; each has `title`, `link`, `description`, `relevance_score`, and `content` when `--include-content`. Full CLI flag list: [`search` skill → `references/flags.md`](../search/references/flags.md).
// JS — see js-sdk-best-practices for all options.
// VERIFIED v1.1.0: discover() returns a WRAPPER { success, data:[...], totalResults, cost, taskId, ... }
const res = await client.discover('Tesla battery tech', { intent: 'EV battery breakthroughs', numResults: 10, includeContent: true });
const rows = res.data; // ← rows are in .data (NOT a bare array, NOT .results)# Python — see python-sdk-best-practices (confirm whether rows come back directly, under .data, or .results) out = client.discover(query="Tesla battery tech", intent="EV battery breakthroughs")
# 1) Trigger
task_id=$(curl -s -X POST https://api.brightdata.com/discover \
-H "Authorization: Bearer $BRIGHTDATA_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"query":"post-quantum cryptography adoption","intent":"enterprise migration guides","mode":"deep","num_results":20,"include_content":true}' \
| jq -r '.task_id')
# 2) Poll until done
while :; do
resp=$(curl -s "https://api.brightdata.com/discover?task_id=$task_id" -H "Authorization: Bearer $BRIGHTDATA_API_TOKEN")
[ "$(echo "$resp" | jq -r '.status')" = "done" ] && break
sleep 3
done
echo "$resp" | jq '.results'`query` is required; everything else is optional. The CLI/SDK expose a subset (see `references/api-reference.md` for the exact per-surface matrix).
| Param | Type | Default | Notes | |---|---|---|---| | `query` | string | — | required, ≤ 1500 chars | | `intent` | string | — | goal descriptor, ≤ 3000 chars; **strongly recommended** — drives ranking | | `mode` | enum | `standard` | `standard` \| `zeroRanking` \| `deep` \| `fast` (REST-only) | | `num_results` | int | — | **1–20**; ignored in `zeroRanking` | | `filter_keywords` | string[] | — | exact keywords that must appear | | `include_content` | bool | `false` | parsed page/PDF content (PDF ≤ 50 MB, 30s); unsupported in `zeroRanking` | | `include_images` | bool | `false` | image array (REST-only) | | `format` | enum | `json` | `json` \| `md` (SDK accepts only `json`) | | `country` | string | `US` | 2-letter ISO | | `city` | string | — | SERP city targeting | | `language` | string | `en` | 31 languages | | `start_date` / `end_date` | string | — | `YYYY-MM-DD` (REST-only) | | `remove_duplicates` | bool | `true` | dedupe results (REST-only) |
| Mode | What it does | Use for | |---|---|---| | `standard` *(default)* | balanced depth + AI ranking | general intent search | | `deep` | exhaustive, broader search; slower | *
Repo: brightdata/skills
Onboard an agent to Bright Data. Use when a coding agent first encounters Bright Data — for…
Social listening and brand reputation research using Bright Data's web scraping…
Debug Bright Data Scraping Browser sessions using the Browser Sessions API. Use this skill…
Build production-ready Bright Data integrations with best practices baked in. Reference…
Bright Data MCP handles ALL web data operations. Replaces WebFetch, WebSearch, and all…
Guide for using the Bright Data CLI (`brightdata` / `bdata`) to scrape websites, search the…