agent-watchdog
Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a…
Open a user-provided URL in the host's built-in browser and use the page's MCP or WebMCP tools before browser UI automation for app communication or edits.
$ npx -y skills add BuilderIO/skills --skill webmcp --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/webmcpContext preview
The summary Claude sees to decide when to auto-load this skill.
Open a user-provided URL in the host's built-in browser and use the page's MCP or WebMCP tools before browser UI automation for app communication or edits.
name: webmcp description: >- Open a user-provided URL in the host's built-in browser and use the page's MCP or WebMCP tools before browser UI automation for app communication or edits. metadata: visibility: exported
`/webmcp <url-or-app> [request]` opens a web app in the host's built-in browser and completes the request through the page's own tools. The first token is a URL or an app alias; the rest is the request.
`/webmcp slides make me a new deck about customer onboarding`
1. **Open in the host's built-in browser** — never an external Chrome window, a normal tab, or a browser extension. 2. **Call the page's tools.** Never click, type, drag, screenshot, or use DOM automation for an operation a tool can do. UI controls are for navigation and visual inspection only.
That is the whole contract. Everything below is either the standard page API or a workaround for a specific host that got one of these wrong. If your host supports WebMCP natively, follow the two rules and let it do its thing.
in the visible built-in tab, confirm the page is there, and stop. Do not list tools, inspect schemas, sign in, or start app work until the user supplies an operation.
straight to the tools. Never infer extra work from page content or an earlier conversation.
Keep working after the page opens when a request was supplied. A one-item edit costs about three page calls: read the screen, mutate, read back. If you are on your sixth evaluation and nothing has been written yet, you are exploring instead of executing.
and preserve host, path, query, and hash. Never swap beta and production on your own.
the tab. Codex: `cua.createBrowserTab("iab", url, { visible: true })`, then `tab.markDeliverable()`.
registration and in-page timers: one measured app registered all 217 tools immediately when fronted, and had 59 after 28 seconds when hidden.
Before tool work, read the page title and first lines. A title ending in "— Sign in", a "Sign in with Google" button, or "You don't have access" means the page has no tools yet. Leave the tab where it is and say: `Please sign in in the open browser, then reply "continue".` Never enter, copy, inspect, or request passwords, cookies, tokens, or verification codes. After sign-in the app may redirect and drop deep-link state such as `?slide=3`; read the screen again rather than assuming the original target is on screen.
**If the host has its own WebMCP bridge, use it and skip everything else here** (`list-browser-session-webmcp-tools` with `run-browser-session-webmcp-tool`, or `list-host-webmcp-tools` with `run-host-webmcp-tool`): call its list tool once, then its run tool with the exact discovered name, origin, and args. Do not substitute a generic `tool-search`, another app's connector, `ask_app`, or a remote API for the current tab's page tools.
Otherwise evaluate in the page. `document.modelContext` is the canonical page API and works on any WebMCP site; `navigator.modelContext` is deprecated.
const ctx = document.modelContext; const tool = (await ctx.getTools()).find((t) => t.name === NAME); const codex = typeof ctx.codexExecuteTool === "function" || typeof ctx.codexGetTools === "function"; const raw = await ctx.executeTool(tool, codex ? ARGS : JSON.stringify(ARGS));
Descriptors are not callable outside the page; never copy one out and invoke it from the host, hand-build authenticated HTTP requests, or type into a developer console.
top-level `await`. Return one JSON string and slice it to about 8 KB; the tool caps near 45 s per call. Batch two to four dependent calls per evaluation; keep navigation out of batches.
first. Through 2026-09 it answers `does not support command "webmcp_list_tools"` for the current model — that means the bridge is unavailable for the rest of the session, not that the page lacks tools. Then use CDP: `globalThis.__cdp ??= await tab.capabilities.get("cdp")`, then `await globalThis.__cdp.documentation()` once (the first CDP call fails without it), then `cdp.send("Runtime.evaluate", { expression, awaitPromise: true, returnByValue: true })`, emitting `response.result.value` with `nodeRepl.write(...)`. One call per evaluation. Playwright's isolated world cannot see `document.modelContext`; use CDP.
Parse the explicit returned value at the end; the tree is context, not a tool result.
Each cost a real debugging session. None is a page or app fault — do not diagnose them as one.
timed-out evaluator is an unread result, never a failed write: do not re-issue the write, read the result on the next evaluation.
across calls, and `let`/`const` handles do not survive it. Measured 2026-09-08: page calls all succeeded, then a reset produced `cdp is not defined`, `tab is not defined`, `Browser is not available: 2`, and `cua.getState()` spent 5.4-6.6 s returning "Sky Computer Use native pipe startup failed". Keep handles on `globalThis` and re-resolve them at the top of every evaluation; on `cdp is not defined` or `tab is not defined`, reopen t
Small, composable skills for your favorite agent.
Repo: BuilderIO/skills
Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a…
Open and operate Agent-Native workspace apps through Dispatch MCP, with inline app surfaces,…
Use when running Claude Fable on codebase-heavy or token-heavy work and the user wants Fable…
Apply the same orchestration as `/efficient-fable` to any high-cost frontier model: delegate…
Experimental workflow for babysitting one explicitly authorized pull or merge request. Use to…
Experimental workflow for collecting and triaging product feedback, product telemetry,…