add-webmcp
Analyze an existing web application, identify safe user-visible capabilities across routes, forms, server actions, handlers, and schemas, then implement…
Builds and improves Browserbase Agent API demos through an Autobrowse-style outer loop: run a fixed task, collect Agent messages and session logs, score the result, revise one system-prompt heuristic, and confirm convergence. Use when creating a Browserbase Agents demo or POC,
$ npx -y skills add browserbase/skills --skill optimize-agent-prompt --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/optimize-agent-promptContext preview
The summary Claude sees to decide when to auto-load this skill.
Builds and improves Browserbase Agent API demos through an Autobrowse-style outer loop: run a fixed task, collect Agent messages and session logs, score the result, revise one system-prompt heuristic, and confirm convergence. Use when creating a Browserbase Agents demo or POC,
name: optimize-agent-prompt description: "Builds and improves Browserbase Agent API demos through an Autobrowse-style outer loop: run a fixed task, collect Agent messages and session logs, score the result, revise one system-prompt heuristic, and confirm convergence. Use when creating a Browserbase Agents demo or POC, optimizing an Agent system prompt, diagnosing flaky Agent runs, or applying auto-research/autobrowse to the Browserbase Agents API." license: MIT allowed-tools: Bash Read Write Edit Grep Glob
Optimize a Browserbase Agent's `systemPrompt` while holding its task, result schema, variables, and evaluation criteria fixed. Treat the outer agent as the teacher and each Browserbase Agent run as an inner-agent rollout.
Use Node.js 18 or later and set `BROWSERBASE_API_KEY`. The harness uses only Node.js built-in modules.
Choose a short experiment name and create an isolated workspace inside the demo or POC repository:
node <skill-dir>/scripts/optimize_agent_prompt.mjs init \ --workspace ./agent-prompt-optimization/<experiment-name> \ --name <experiment-name>
Edit the generated files:
Use concrete success criteria. Prefer a strict JSON Schema with required fields and `null` for unavailable facts. Add known-field regexes and factuality-warning regexes under `evaluation` when a truth oracle exists. Read [references/evaluation.md](references/evaluation.md) when designing the task or score.
node <skill-dir>/scripts/optimize_agent_prompt.mjs run \ --workspace ./agent-prompt-optimization/<experiment-name> \ --prompt prompts/iteration-001.md \ --label iteration-001
The harness creates one reusable Browserbase Agent, updates its `systemPrompt` on later iterations, starts the run, polls messages and status, and writes:
runs/<label>/ ├── system-prompt.md ├── created-run.json ├── run.json ├── messages.json ├── session-logs.json └── summary.json
It stops a run after the configured message budget instead of paying for an unproductive spiral. Use `--max-messages`, `--timeout-ms`, `--proxies`, or `--verified` only when the task needs different values from `task.json`.
Start with the compact trajectory:
node <skill-dir>/scripts/optimize_agent_prompt.mjs inspect \ --workspace ./agent-prompt-optimization/<experiment-name> \ --label iteration-001
Then read `summary.json` and drill into `messages.json` at the first wrong or wasted turn. Agent messages expose ordered tool calls, tool results, errors, and final output. A `reasoning` part may contain no readable text; never require hidden chain-of-thought for the teacher loop.
Read `session-logs.json` only when browser-level evidence can distinguish the cause—for example, a redirect, 403, failed request, console error, or hidden endpoint. Empty session logs can mean the Agent completed with search/fetch tools and never drove its browser.
See [references/api.md](references/api.md) for endpoint shapes, pagination, result normalization, and trace caveats.
Find the earliest consequential failure and state one counterfactual:
> If the system prompt had instructed X, the Agent would have avoided Y, as shown by tool result Z.
Copy the current prompt to `prompts/iteration-NNN.md` and make one attributable change. Typical improvements are:
Keep wins. If the new run regresses, restore the previous prompt and test a different hypothesis rather than stacking more rules.
Generate the comparison table after each run:
node <skill-dir>/scripts/optimize_agent_prompt.mjs report \ --workspace ./agent-prompt-optimization/<experiment-name>
Judge more than field completeness. Require:
Once a prompt wins, run it again unchanged with a new label. Converge only after it passes at least two of the last three runs and one pass is an unchanged confirmation. Do not call a prompt globally optimal from one task; describe it as the best prompt for the tested task distribution.
Use the confirmed prompt as the Agent's production `systemPrompt`. Keep the strict result schema and per-run variables. Preserve the experiment workspace or its report so reviewers can audit why each instruction exists.
In the final handoff, report:
A set of skills for enabling Claude Code to work with Browserbase through browser automation and the official browse CLI.
Repo: browserbase/skills
Analyze an existing web application, identify safe user-visible capabilities across routes, forms, server actions, handlers, and schemas, then implement…
Audit the developer experience of a product, SDK, docs site, or SKILL.md by dropping multiple Claude subagents at it with only a tiny task prompt and real…
Self-improving browser automation via the auto-research loop. Iteratively runs a browsing task, reads the trace, and improves the navigation skill…
Turn a website's observable HTTP traffic into a best-effort OpenAPI 3.1 spec by analyzing a `browser-trace` capture. Use when the user wants to…
Capture a full DevTools-protocol trace of any browser automation — CDP firehose, screenshots, and DOM dumps — then bisect the stream into per-page searchable…
Migrate browser-use (Python) browser-automation scripts to Stagehand v3 (TypeScript) on Browserbase. Use when the user wants to convert, port, rewrite, or…