cua-driver
Required host-generated bearer token when the local MCP HTTP endpoint is enabled.
Build or adapt a bounded computer-use loop where Cua Driver observes and acts, TypeSafe Jev selects only from application-owned candidate IDs, and the caller validates and verifies every action. Use for the jev-use recipe or similar Jev integrations; do not use it to add model
$ npx -y skills add trycua/cua --skill jev-use --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/jev-useContext preview
The summary Claude sees to decide when to auto-load this skill.
Build or adapt a bounded computer-use loop where Cua Driver observes and acts, TypeSafe Jev selects only from application-owned candidate IDs, and the caller validates and verifies every action. Use for the jev-use recipe or similar Jev integrations; do not use it to add model
name: jev-use description: Build or adapt a bounded computer-use loop where Cua Driver observes and acts, TypeSafe Jev selects only from application-owned candidate IDs, and the caller validates and verifies every action. Use for the jev-use recipe or similar Jev integrations; do not use it to add model logic or credentials to Cua Driver.
Keep the decision layer above Cua Driver. Driver supplies observations and executes actions; the application constructs complete candidates; TypeSafe Jev returns one candidate ID. Never let Jev invent tool names, coordinates, refs, targets, delivery modes, or other arguments.
Use the example at `libs/cua-driver/examples/jev-use/` as the runnable reference. Keep TypeSafe request construction in the external Jev adapter rather than in Driver or a Driver extension. The Python and TypeScript adapters must expose equivalent mock and live behavior. For a process boundary, use `cua.jev_choice_request_v1` on stdin and require `cua.jev_choice_v1` on stdout. The request contains only a goal, capture ID, compact regions, bounded history, and candidate IDs with descriptions; the response contains only the selected ID, model identity, confidence, and probabilities. Invoke the Python interpreter and absolute chooser path directly without a shell. For native desktop applications, use `NativeAccessibilitySource` and `cua.jev_choice_request_v2`, which adds a per-candidate `source` (`page`, `ax`, or `visual`), compact value-free `elements`, and optional `progress` counted from the runner's own performed actions. Browser tasks keep sending v1. Prefer browser DOM and semantic evidence. The optional visual adapter consumes the public `cua.visual_regions_v1` result only when Driver advertises both `parse_visual_regions` and the capture-bound `click.capture_id` input. Use the checked-in fixtures for deterministic development; do not add a model, extension artifact, or Driver implementation detail to the recipe.
1. State the goal and obtain a fresh Cua Driver observation through one persistent CLI or MCP session. 2. Prefer an unambiguous fresh accessibility or browser DOM token. 3. If visual grounding is needed, discover `parse_visual_regions` through the current MCP tool inventory. Validate its versioned result, capture ID, screenshot reference and dimensions, coordinate mapping, unique region IDs, bounds, content, confidence, and ambiguity. Build a pixel action only with the exact capture ID in the same `click` call. Otherwise reobserve or abstain. 4. Construct a bounded candidate table. Each executable candidate contains the complete Driver tool and arguments. Include `reobserve` and `abstain` when evidence can be stale, incomplete, or ambiguous. 5. Send Jev only the goal, compact observation, recent history, and candidate IDs with descriptions. Include typed visual regions and their `capture_id` when the current observation has validated visual evidence; do not send extension internals or screenshot bytes. 6. Resolve the returned ID against the original immutable table. Reject an unknown, duplicate, malformed, denied, stale, or capture-mismatched choice, or a result below the caller's stated confidence policy. 7. Execute at most one Driver action. Use background delivery by default; foreground delivery is an explicit escalation subject to the active Driver contract and user authorization. 8. Reobserve and verify the postcondition before building another table.
region IDs as observation-local. Never reuse them after the UI changes.
coordinate space and tied to the same target and snapshot.
offer `reobserve` and `abstain` without inventing a mutation.
an unbound coordinate action.
not prove editability or interactivity.
tree and the screenshot together, so element tokens and `capture_id` describe the same moment.
`native_roles.py` / `native_roles.ts`, keyed by Driver's `normalized_role`. Do not normalize roles in Driver.
`in_web_content` elements, window chrome, and labels equal to the value.
never `element_index`. Cap at 24 action candidates plus `reobserve` and `abstain`, and log how many were dropped.
allows that risk. Text comes only from task parameters.
unbound coordinate action.
harness task-state file, not the accessibility tree the model saw.
The deterministic mock path must work without `TYPESAFE_API_KEY`. For live Jev, read the key from the process environment or a secure interactive prompt; never put it in source, command arguments, logs, artifacts, or messages. Verify task completion from an independent application postcondition rather than a model answer, action response, or screenshot alone.
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
Repo: trycua/cua
Required host-generated bearer token when the local MCP HTTP endpoint is enabled.
Create, use and clean up cua sandboxes (disposable Linux or macOS computers) locally or in…
Work inside cua Spaces through the cua MCP server. A Space is a remote or local computer the…
Use Cua Volume, the one volume every Space and agent of this user shares. Inside a Space it…
Use when you need to visually interact with a GUI: test buttons, fill forms, verify visual…