bug-reproduce
Turn a known bug into a tight, red-capable reproducer, then prove the reproducer locks that…
Drive the desktop background-first; escalate on signal.
$ npx -y skills add Prismer-AI/PrismerCloud --skill prismer-computer-use --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/prismer-computer-useContext preview
The summary Claude sees to decide when to auto-load this skill.
Drive the desktop background-first; escalate on signal.
name: prismer-computer-use
scope: common
description: "Drive the desktop background-first; escalate on signal."
version: 2.1.0
author: Francesco Bonacci (f-trycua), Hermes Agent
license: MIT
platforms: [ macos, windows, linux ]
metadata:
nativeReplaces: [ computer-use ]
hermes:
tags: [ computer-use, desktop, automation, gui, cross-platform ]
category: desktop
related_skills: []
requiresExplicitGrant: trueThis is the uniquely named `prismer-computer-use` skill, adapted from Hermes. Use the actual tools exposed by the executing host; examples using terminal, process, delegate_task, vision_analyze or browser_* are not tool registrations. Missing dependencies do not hide this skill. Report command startup, version, account/permissions and task-specific live verification separately. Use task-owned artifact paths and existing user authorization; do not change shared accounts, Runtime/provider configuration, global security settings or unrelated work. See NOTICE.md and LICENSE for resource provenance. Runtime availability and upstream-entry suppression are owned by the integration layer.
You have a `computer_use` tool that drives the user's desktop in the **background** — your actions do NOT move the user's cursor, steal keyboard focus, or switch virtual desktops / Spaces. The user can keep typing in their editor while you click around in a browser in another window. This is the opposite of pyautogui-style automation.
Everything here works with any tool-capable model — Claude, GPT, Gemini, or an open model on a local OpenAI-compatible endpoint. There is no Anthropic-native schema to learn.
Hermes drives [cua-driver](https://github.com/trycua/cua) under the hood. This skill teaches the Hermes `computer_use` **action vocabulary**, which is NOT the driver's raw MCP vocabulary. Call the actions documented below and never the driver's tools by name: `capture` is a Hermes action that maps to the driver's `get_window_state`; `element=N` is a Hermes argument that the wrapper translates into the driver's `element_token` handle. If you see a driver-side error mentioning `snapshot_id`, `element_token`, or "no reviewed risk classification", you (or a stale description) called the raw driver vocabulary — go back to the actions below.
**Step 1 — Capture first.** Almost every task starts with:
computer_use(action="capture", mode="som", app="<the app you're driving>")
Returns a screenshot plus an indexed element list like:
#1 AXButton 'Back' @ (12, 80, 28, 28) [Chrome] #2 AXTextField 'Address bar' @ (80, 80, 900, 32) [Chrome] #7 Link 'Sign In' @ (900, 420, 80, 24) [Chrome] ...
The `#N` index is the ONLY element handle you use. Behind it the wrapper keeps this snapshot's opaque per-element token and sends it with every `element=N` action, so a click on an index from a superseded snapshot is refused explicitly (`stale`) instead of landing on the wrong control. Re-capture after anything that changes the screen; indices do not survive it.
The role names match the host platform's accessibility framework (`AXButton` on macOS, `Button` on Windows UIA, `push button` on Linux AT-SPI) — treat them as labels, not as strict types.
**Step 2 — Click by element index.** This is the single most important habit:
computer_use(action="click", element=7)
Much more reliable than pixel coordinates for every model. Claude was trained on both; other models are often only reliable with indices.
**Step 3 — Verify.** After any state-changing action, re-capture. You can save a round-trip by asking for the post-action capture inline:
computer_use(action="click", element=7, capture_after=True)
| `mode` | Returns | Best for | |---|---|---| | `som` (default) | Screenshot + indexed element list | Vision models; preferred default | | `vision` | Plain screenshot, no elements | When you only need pixels (then click by `coordinate=`) | | `ax` | Element list only, no image | Text-only models, or when you don't need to see pixels |
Current drivers always return the screenshot AND the tree in one call; `mode` decides what Hermes hands back to you, not what the driver does. There is no numbered overlay burned into the screenshot — the index list is the map; ground on both and cross-check (the tree lies on some surfaces).
**No vision model?** If your main model can't read images (or the provider rejects image tool results), Hermes routes the screenshot through the auxiliary vision model and you get a text description instead of pixels. Configure `auxiliary.vision` in `config.yaml` to pick that model, or use `mode="ax"` and drive by element index without a screenshot at all.
capture mode=som|vision|ax app=… (default: current app) click element=N OR coordinate=[x, y] button=left|right|middle double_click element=N OR coordinate=[x, y] right_click element=N OR coordinate=[x, y] middle_click element=N OR coordinate=[x, y] drag from_element=N, to_element=M (or from/to_coordinate) scroll direction=up|down|left|right amount=3 (ticks) type text="…" key keys="<save shortcut>" | "return" | "escape" | "<modifier>+t" set_value element=N value="…" (selects/sliders without opening the menu) wait seconds=0.5 list_apps list_windows focus_app app="<app name>" raise_window=false (default: don't raise)
All actions accept optional `capture_after=True` to get a follow-up screenshot in the same tool call. All actions that target an element accept `modifiers=[…]` for held keys.
The input actions (`click`, `double_click`, `right_click`, `middle_click`, `drag`, `scroll`, `type`, `key`) also accept `delivery_mode`. The optional `bring_to_front=True` request invokes a separate
Repo: Prismer-AI/PrismerCloud
Turn a known bug into a tight, red-capable reproducer, then prove the reproducer locks that…
Review a diff against its acceptance criteria in four segments (convention adherence, bug…
Five-dimension design audit (frontend UI/UX · server data-model & flow · endpoint spec ·…
Before merge, mechanize Documentation-First — derive the code delta from git diff, then…
Diagnose the local dev machine before any APC loop step — run apc env doctor, classify each…
Close out a local coding task on the bound daemon — stage, commit, branch, merge, push via…