Skip to content
Development
Skill

/computer-control

Automates desktop GUI workflows via computer use API with screenshot capture. Use when scripting GUI interactions or recording browser sessions for tutorials.

From plugin
claude-night-market
337200 skills59 agents162 commands1 MCP
Install
$ npx -y skills add athola/claude-night-market --skill computer-control --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/computer-control

Context preview

The summary Claude sees to decide when to auto-load this skill.

Automates desktop GUI workflows via computer use API with screenshot capture. Use when scripting GUI interactions or recording browser sessions for tutorials.

SKILL.md

computer-control.SKILL.md
name: computer-control
description: Automates desktop GUI workflows via computer use API with screenshot capture. Use when scripting GUI interactions or recording browser sessions for tutorials.
alwaysApply: false
model_hint: standard

Computer Control Skill

Use Claude's Computer Use API to see and control desktop environments through screenshots and mouse/keyboard actions.

When To Use

  • Automating GUI-based workflows that lack CLI alternatives
  • Testing web applications through visual interaction
  • Filling forms, navigating menus, or interacting with desktop apps
  • Building automation pipelines that need visual verification

When NOT To Use

  • Tasks achievable through CLI or API (no GUI needed)
  • Browser automation better served by Playwright or CDP

> **Why this stays opt-in.** Per > [docs/inclusive-defaults.md][inc] (TRUE-exception > category 4), Computer Use takes screenshots and > synthesizes keyboard/mouse input: cross-process side > effects that must always be explicitly invoked, never > default-on.

[inc]: ../../../../docs/inclusive-defaults.md

Architecture

The computer use system has three layers:

1. **Display Toolkit** (`phantom.display`) - executes OS-level actions via xdotool/scrot on the real or virtual display 2. **Agent Loop** (`phantom.loop`) - manages the conversation cycle between Claude API and the display toolkit 3. **CLI** (`phantom.cli`) - command-line interface for running tasks or checking environment readiness

User Task
    |
    v
Agent Loop  <---->  Claude API (beta)
    |                   |
    v                   v
Display Toolkit    tool_use responses
    |              (click, type, screenshot)
    v
OS Commands (xdotool, scrot)
    |
    v
Display (X11 / Xvfb / WSLg)

Quick Start

Check environment

cd plugins/phantom
uv run python -m phantom.cli --check

Run a task

export ANTHROPIC_API_KEY="sk-ant-..."
uv run python -m phantom.cli "Open Firefox and search for Claude AI"

Use in Python

from phantom.display import DisplayConfig, DisplayToolkit
from phantom.loop import LoopConfig, run_loop

result = run_loop(
    task="Take a screenshot of the desktop",
    api_key="sk-ant-...",
    loop_config=LoopConfig(
        model="claude-sonnet-5",
        max_iterations=10,
    ),
    display_config=DisplayConfig(width=1920, height=1080),
)

print(f"Done in {result.iterations} iterations")
print(result.final_text)

API Versions

| Model | Tool Version | Beta Flag | |-------|-------------|-----------| | Opus 4.6, Sonnet 4.6, Opus 4.5 | `computer_20251124` | `computer-use-2025-11-24` | | Sonnet 4.5, Haiku 4.5, older | `computer_20250124` | `computer-use-2025-01-24` |

The `resolve_tool_version()` function handles this mapping automatically based on the model name.

Available Actions

**All versions:**

  • `screenshot` - capture display
  • `left_click` - click at `[x, y]`
  • `type` - type text string
  • `key` - press key combo (e.g., `ctrl+s`)
  • `mouse_move` - move cursor

**Enhanced (20250124+):**

  • `scroll` - scroll with direction and amount
  • `left_click_drag` - drag between coordinates
  • `right_click`, `middle_click`, `double_click`, `triple_click`
  • `hold_key` - hold key for duration
  • `wait` - pause between actions

**Latest (20251124):**

  • `zoom` - inspect screen region at full resolution

Safety

Computer use carries risks. Follow these guidelines:

1. **Use a sandbox**: Run in Docker or a VM, not your main OS 2. **Limit access**: Do not provide login credentials unless necessary, and never for banking or sensitive services 3. **Set iteration caps**: Always use `max_iterations` to prevent runaway API costs 4. **Human approval**: For actions with real-world consequences, add confirmation callbacks via `on_action` 5. **Close sensitive apps**: Claude sees the full screen via screenshots; close anything private before starting

Environment Requirements

**Linux (native or WSL2 with WSLg):**

sudo apt install xdotool scrot xclip

**Headless (Docker/CI):**

# Install Xvfb for virtual display
sudo apt install xvfb xdotool scrot xclip
Xvfb :1 -screen 0 1920x1080x24 &
export DISPLAY=:1

Prompting Tips

1. Be specific about each step of the task 2. Add "After each step, take a screenshot and verify" to catch mistakes early 3. Use keyboard shortcuts when UI elements are hard to click 4. Provide example screenshots for repeatable workflows 5. Set a system prompt with domain-specific instructions

Exit Criteria

  • [ ] `uv run python -m phantom.cli --check` exits 0 before any

task is launched; if it fails, required OS tools (xdotool, scrot, xclip) are installed or Xvfb is started before proceeding

  • [ ] `max_iterations` set on every `run_loop()` call; no task

launched without an explicit iteration cap to prevent runaway API costs

  • [ ] Final result includes `result.iterations` count and

`result.final_text` confirming the task outcome; empty `final_text` treated as failure, not success

  • [ ] Sensitive applications (password managers, banking, private

files) closed before task starts; task prompt does not contain raw credentials

Read more
Ships withclaude-night-market

A plugin marketplace for Claude Code. Install only the plugins you need to run git workflows, code review, spec-driven development, and autonomous agents from inside your Claude Code session.

Get the whole plugin

Other skills on claude-night-market.