Skip to content

/agent-browser

Browser automation CLI for AI agents. Use when the user needs to inspect, test, or automate browser behavior: navigating pages, filling forms, clicking buttons, taking screenshots, extracting page data, reading selected OpenDesign browser-tab context, testing web apps,

From plugin
open-design
96k164 skills1 command1 MCP
Install
$ npx -y skills add nexu-io/open-design --skill agent-browser --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/agent-browser

Context preview

The summary Claude sees to decide when to auto-load this skill.

Browser automation CLI for AI agents. Use when the user needs to inspect, test, or automate browser behavior: navigating pages, filling forms, clicking buttons, taking screenshots, extracting page data, reading selected OpenDesign browser-tab context, testing web apps,

SKILL.md

agent-browser.SKILL.md
name: agent-browser
description: |
  Browser automation CLI for AI agents. Use when the user needs to inspect,
  test, or automate browser behavior: navigating pages, filling forms,
  clicking buttons, taking screenshots, extracting page data, reading selected
  OpenDesign browser-tab context, testing web apps, dogfooding OpenDesign
  previews, QA, bug hunts, or reviewing app quality. Prefer local OpenDesign
  preview URLs unless the user explicitly asks for external browsing.
triggers:
  - "browser"
  - "current browser tab"
  - "selected tab"
  - "open website"
  - "test this web app"
  - "take a screenshot"
  - "element screenshot"
  - "extract logo"
  - "extract fonts"
  - "extract colors"
  - "extract images"
  - "extract motion"
  - "OG metadata"
  - "accessibility"
  - "a11y"
  - "click a button"
  - "fill out a form"
  - "scrape page"
  - "QA"
  - "dogfood"
  - "bug hunt"
od:
  mode: prototype
  surface: web
  platform: desktop
  scenario: validation
  preview:
    type: markdown
  design_system:
    requires: false
  upstream: "https://github.com/vercel-labs/agent-browser/blob/main/skills/agent-browser/SKILL.md"
  capabilities_required:
    - file_write

Agent Browser

Use `agent-browser` for local OpenDesign preview validation: inspect rendered state, click/type when requested, and capture one screenshot when visual evidence matters. Keep the browser local-first unless the user explicitly asks for external browsing.

When the run prompt contains selected workspace context, prefer the selected `browser` tab URL/title as the target. Treat user phrases like "this page", "the current browser", "right-side tab", "extract the logo", "get the palette", "take an element screenshot", or "check OG/a11y" as requests about that selected tab unless the user names another target.

Requirements

Verify the CLI before doing any browser work:

command -v agent-browser

If missing, stop and tell the user to install it:

npm i -g agent-browser
agent-browser install

Do not replace the CLI with ad hoc browser scripts.

Context Hygiene

Never print full upstream guides into chat or tool output. Save them to temp files and extract only task-relevant lines:

AGENT_BROWSER_CORE="${TMPDIR:-/tmp}/agent-browser-core.$$.md"
agent-browser skills get core > "$AGENT_BROWSER_CORE"
rg -n "cdp|connect|snapshot|screenshot|click|type|wait|get title|get url" "$AGENT_BROWSER_CORE"

Use `agent-browser skills get core --full` only when needed, and redirect it to a temp file the same way.

Browser Context Extraction

For selected OpenDesign browser tabs and browser-use/browser-harness-style tasks, collect the smallest useful evidence first:

1. Confirm the target with `agent-browser get title` and `agent-browser get url`. 2. Capture `agent-browser snapshot` before any extraction or click. 3. For visual evidence, save a page screenshot and, when the core guide exposes an element-screenshot command, capture the specific element instead of a cropped full page. 4. For logos, fonts, colors, images, motion code, OG metadata, page structure, and accessibility checks, prefer DOM/CSS/accessibility evidence from the attached browser over guessing from the rendered screenshot alone. 5. If the selected OpenDesign context only provided a URL/title and no browser automation tool is attached, say that directly and do not invent page internals.

Save extracted design evidence as compact notes or assets in the project when the user is building from the reference. Do not paste full page HTML or large asset dumps into chat; summarize the relevant selectors, tokens, URLs, and screenshots.

CDP Startup Contract

`agent-browser` must attach to an existing CDP endpoint. Never run `agent-browser open` before `agent-browser connect`; doing so can make the CLI auto-launch Chrome and re-enter the crash path.

Do not run OpenDesign's own daemon CLI as a browser automation tool. Commands such as `od browser snapshot`, `daemon-cli.mjs browser snapshot`, or `$OD_NODE_BIN $OD_BIN browser snapshot` are not valid browser tools; they can be misinterpreted as daemon startup and open an internal `127.0.0.1:<port>` service in the system browser. Use the external `agent-browser` CLI attached to CDP instead.

Use this sequence:

if ! curl -fsS http://127.0.0.1:9223/json/version | rg -q webSocketDebuggerUrl; then
  open -na "Google Chrome" --args \
    --remote-debugging-port=9223 \
    --user-data-dir=/tmp/od-agent-browser-chrome \
    --no-first-run \
    --no-default-browser-check

  for i in {1..20}; do
    if curl -fsS http://127.0.0.1:9223/json/version | rg -q webSocketDebuggerUrl; then
      break
    fi
    sleep 0.5
  done
fi

curl -fsS http://127.0.0.1:9223/json/version | rg webSocketDebuggerUrl
agent-browser connect http://127.0.0.1:9223

If CDP is still unavailable after polling, stop and ask the user to launch Chrome manually from Terminal:

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \
  --remote-debugging-port=9223 \
  --user-data-dir=/tmp/od-agent-browser-chrome \
  --no-first-run \
  --no-default-browser-check

If Chrome exits before CDP is ready or reports `DevToolsActivePort`, report: "Chrome crashed before CDP became available; start Chrome manually with `--remote-debugging-port` and retry attach."

Lightpanda is optional. Do not try `--engine lightpanda` unless `command -v lightpanda` succeeds.

OpenDesign Smoke Path

Use a temp home and stable session:

export HOME=/tmp/agent-browser-home
export AGENT_BROWSER_SESSION=od-local-preview

When you start a temporary Chrome profile for this smoke path, close it before finishing the task. Prefer a shell trap around the whole smoke script:

CHROME_USER_DATA_DIR=/tmp/od-agent-browser-chrome
cleanup_agent_browser() {
  pkill -f -- "--user-data-dir=${CHROME_USER_DATA_DIR}" 2>/dev/null || true
}
trap cleanup_agent_browser EXIT INT TERM

With the OpenDesig

Read more
Ships withopen-design

🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.

Get the whole plugin, auto-invoked

Other skills on open-design.