Skip to content
Agent Orchestration
Skill

/browser-automation

Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

From plugin
openclaw
390k68 skills
Install
$ npx -y skills add steipete/clawdis --skill browser-automation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/browser-automation

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

SKILL.md

browser-automation.SKILL.md
name: browser-automation
description: Use when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.
user-invocable: false

Browser Automation

Use this skill when you need the `browser` tool for anything beyond a single page check.

Operating Loop

1. Check browser state before acting:

  • `openclaw browser doctor` or `action="status"` when the browser/plugin setup itself may be broken.
  • `action="status"` for availability.
  • `action="profiles"` if login state or profile choice matters.
  • `action="tabs"` before opening a new tab if retries/timeouts may have left windows behind.

2. Prefer stable tab handles:

  • Open important tabs with `label`, for example `label="meet"`.
  • After `action="tabs"` or `action="open"`, store `suggestedTargetId` and pass it as `targetId` in later calls.
  • `suggestedTargetId` is the label when one exists, otherwise the stable `tabId` handle like `t1`.
  • Avoid relying on raw DevTools `targetId` except for immediate diagnostics; it can change under Chromium target replacement.

3. Read before you click:

  • For “read the page and answer X,” use `action="text"` with optional `selector` and `maxChars` for bounded visible prose (first selector match, otherwise article/main/body). On existing-session profiles, use `snapshot` instead. Efficient snapshots omit most prose.
  • For virtualized lists, scroll through each segment, capture only the relevant rows, then merge the results.
  • Use `action="snapshot"` on the intended `targetId`.
  • Add snapshot `query` to find lines containing all query tokens, ignoring case; matching lines keep their refs.
  • Use the same `targetId` for follow-up actions so refs stay on the same tab.
  • For durable Playwright refs, request `refs="aria"` when supported. If you receive `axN` refs from `snapshotFormat="aria"`, use them only after that same snapshot call; stale or unbound `axN` refs fail fast and need a fresh snapshot.
  • Use `urls=true` when link text is ambiguous or a direct navigation target would avoid brittle clicks.
  • Use `labels=true` on snapshot or screenshot when visual position matters. On Playwright-backed profiles, the response includes an `annotations` array (`{ref, number, role, name?, box}`) with each ref's bounding box in the captured image's coordinate space, so you can reason about position without re-snapshotting; screenshot labels can also combine with `fullPage=true` (CLI: `--full-page`) to label the whole document, or `ref` / `element` to clip to one element. `profile="user"` and other existing-session (chrome-mcp) profiles render an overlay into page screenshots but do not attach `annotations` or use the Playwright full-page/ref/element projection helper, so read positions from the labeled image itself on those profiles. The raw-CDP fallback (no Playwright) does not support labeled screenshots at all and returns a 501, so only request `labels` when Playwright is available.

4. Act narrowly:

  • Prefer `action="act"` with a ref from the latest snapshot.
  • `navigate` returns the loaded page's compact snapshot inline, and batch `act` results that report a cross-document navigation include fresh page state; use those refs directly instead of a follow-up snapshot call.
  • After a single act that triggers navigation, and after modal changes or form submissions, snapshot again before the next action.
  • Avoid blind waits. Wait for visible UI state when possible.
  • Use `action="emulate"` with `device`, `colorScheme`, `timezoneId`, or `locale` when testing those settings; snapshot again afterward. Existing-session profiles do not support emulation.

5. Report real blockers:

  • Debug network failures with `action="requests"`, optional URL/type `filter`, and `limit` (default 50 recent entries). `clear=true` clears the collected log after reading. Use a managed profile; existing-session profiles do not support this log.
  • Debug page errors with `action="errors"` and `limit` (default 50 recent entries). `clear=true` clears the collected log after reading. Existing-session profiles do not support this log.
  • If the page needs login, permission, captcha, 2FA, camera/microphone approval, or another manual step, stop and tell the user exactly what is needed.
  • Do not claim the browser is not logged in just because the current page shows a permission or onboarding dialog. Inspect the visible UI first.

Browser batch CLI

`openclaw browser batch` runs an array of nested `/act` actions in one `/act` call (the same `kind="batch"` runtime reached through the agent tool), so CLI users and scripts can combine actions like `wait`, `click`, `type`, and `evaluate` into a single replayable plan without per-action round trips. Each entry in `actions[]` is a `BrowserActRequest` — the closed union the `/act` route accepts — not arbitrary `openclaw browser` subcommands. `batch` is not supported on `profile="user"` and other existing-session (chrome-mcp) profiles; send actions individually there.

  • CLI: `openclaw browser batch --actions '<json>'`, `--actions-file plan.json`, or `--actions-file -` for stdin. `--actions-file` and stdin input are capped at 1,000,000 bytes; split larger plans into multiple batch commands. `--continue` sets `stopOnError=false`; default stops on first error.
  • Ref lifecycle: refs come from a `snapshot` run before the batch (snapshot is not a nested action). A nested action that changes page state — such as a `click` that triggers navigation, or an `evaluate` that mutates the DOM — can invalidate earlier refs for the rest of the batch; put state-changing actions first, or split into a follow-up batch after re-snapshotting. Navigation and re-snapshotting happen outside the batch, since `open`, `navigate`, and `snapshot` are not `/act` kinds.
  • Target id: nested actions share the request's tab; an explicit nested `targetId` that resolves to a differe
Read more
Ships withopenclaw

OpenClaw is an open-source AI assistant that runs on your own computer and meets you in the channels you already use: Discord, iMessage, Slack, Teams, Telegram, WhatsApp, and 20+ more, plus native apps for macOS, iOS, Android, Windows, and Linux.

Get the whole plugin
Stats
389,501
Stars
81,881
Forks
Active
Maintenance
TypeScript
Language
1d ago
Last commit
9mo ago
Created
13d ago
Added

Repo: steipete/clawdis

Other skills on openclaw.