an
Open and operate Agent-Native workspace apps through Dispatch MCP, with inline app surfaces,…
Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user
$ npx -y skills add BuilderIO/skills --skill agent-watchdog --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/agent-watchdogContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user
name: agent-watchdog description: Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user asked, independently investigate the same problem to form your own hypotheses and approach, inspect what the agent actually changed and verified, compare the two investigations, report gaps and add-on directions, and optionally make scoped fixes when the user authorizes repair.
Watch another agent's work like a reviewer with a pager who is also a second investigator: wait for completion when needed, reconstruct the request, run your own independent investigation of the same problem, verify the evidence, and close the gap between what was asked and what actually happened. You are not just grading their homework — you are a second solver whose findings get diffed against theirs so nothing is missed.
Infer the mode from the user's wording:
reaches a terminal state. Do not edit files.
or final claims, run the independent investigation below, and return a gap report plus add-on directions. Do not edit files.
broad rewrites, branch movement, or speculative changes.
the same original request and reconcile the important differences.
If authority is unclear, default to audit-only and say what you would fix.
1. Identify every artifact the user supplied: session ID, transcript path, thread URL, PR, branch, commit, CI run, issue, Slack link, or pasted summary. 2. Use the host's native thread/history tools, local transcript files, repo logs, GitHub tools, or pasted content to resolve the artifact. Prefer the most direct source over summaries. 3. If the artifact is still running and the user asked to watch, poll at a reasonable interval until it is done, blocked, stale, or clearly waiting on a human/external system. 4. If the artifact cannot be resolved, ask for the missing identifier or path.
Build a compact contract before judging the work:
versions, validation expectations, design requirements, or security/privacy limits.
screenshots, review replies, or status updates.
Treat the user's request as the source of truth, not the other agent's summary.
Act as if the original prompt had been given to you, in parallel with auditing the watched agent. Anchoring is the failure mode: reading their work first and nodding along. Your value comes from a genuinely separate second pass.
1. Form your own hypotheses about root causes and the approach you would take — ideally before reading the watched agent's conclusions, and if you have already seen them, still reason from first principles rather than from their framing. 2. Explore the code, data, logs, production state, and docs yourself, directly or via subagents. While the watched agent is still running, use the wait to pre-map the problem domains so hypotheses are ready before their diff lands. 3. Prioritize evidence the watched agent may not have looked at: production run ledgers or databases, session replays, error trackers, user-supplied screenshots, deploy/version state, other worktrees, memory of past incidents in the same area. 4. Verify the watched agent's key claims against primary sources, and verify your own leads the same way — reopen the cited files and line refs before asserting either side is right. Subagent reports are leads, not facts. 5. Diff the two investigations: what they found that you missed, what you found that they missed, where the approaches diverge, and any product or design fork they took silently. Convert the diff into concrete, actionable add-on directions (exact files, guards, test cases) — not vague concerns — and offer them as a paste-ready note the user can relay.
Finding something is not a reason to send it yet. A watched agent that is still building re-plans around every note it receives, so the cost of a relay is their attention and their sequencing, not your tokens.
code they are touching now, a correction to something you told them earlier, or an answer they are blocked on.
waits for their next checkpoint and travels as one batched note.
labelled as a map rather than a to-do list, or not at all: an unranked backlog reliably produces several half-finished surfaces instead of a few complete ones.
re-litigating settled ground and keeps the relationship peer-to-peer.
another one.
Inspect evidence, not vibes:
Small, composable skills for your favorite agent.
Repo: BuilderIO/skills
Open and operate Agent-Native workspace apps through Dispatch MCP, with inline app surfaces,…
Use when running Claude Fable on codebase-heavy or token-heavy work and the user wants Fable…
Apply the same orchestration as `/efficient-fable` to any high-cost frontier model: delegate…
Experimental workflow for babysitting one explicitly authorized pull or merge request. Use to…
Experimental workflow for collecting and triaging product feedback, product telemetry,…
Experimental workflow for summarizing work that still needs human judgment across configured…