Library · Automation: browser, screen and schedules

Computer Use: controlling the screen

Builder70 minUpdated: October 2026
40 of 105 in the library

Module: 9. Advanced features | Time: ~30 min theory + 40 min practice


The gist

Think of a remote desktop where the operator is an AI. Claude sees your screen (through screenshots) and can click, type and scroll, but only in the apps you have explicitly allowed (permission: the right to perform an action). This isn't browser automation, and it isn't an API (application programming interface) integration. It's control of native apps: Finder, Calendar, Notes, System Settings, any desktop tool that has no API.

In Claude Code this works through the built-in MCP server computer-use (Computer Use: the mode for controlling your desktop). As of October 2026, in Claude Code for the terminal it's a research preview (an experimental feature) on macOS for the Pro and Max plans; Team and Enterprise don't have it. In the Claude desktop app, the same capability is available on macOS and Windows. At the Anthropic API level, the tool has a current version, computer_toolset_20260801 (no beta header), and an earlier beta version, computer_20251124, for older models.

Terms in this lesson: Computer Use (the mode for controlling your desktop), API (application programming interface), permission (the right to perform an action), prompt (a request to the AI), token (a unit of text for the AI).


Key concepts

  • Computer Use: Claude's ability to control the screen through the MCP server computer-use (in Claude Code) or through the API tool type computer_toolset_20260801 (in the Anthropic API)
  • Tiered access: a three-level system of access to apps (view only / clicks only / full)
  • request_access: an explicit request for permission to control a specific app before any action
  • Screenshot → analyze → act loop: the working cycle (agentic loop): screenshot, analysis, action, check
  • computer_batch: batched commands for several actions in a row without extra round trips
  • Teach mode: an interactive teaching mode with pop-up hints (tooltips) over each action
  • zoom: enlarging part of the screen to read small text (on by default in the current version of the tool)

Theory

Two ways to use it

Computer Use exists on two levels:

1. In Claude Code (through MCP): the built-in computer-use server with a set of tools (screenshot, click, typing and others). It runs locally on your Mac. It's off by default. To turn it on: in an interactive session, type /mcp, find computer-use and choose Enable (the setting is remembered for the project). The first time you use it, macOS will ask for two permissions: Accessibility (for clicks and typing) and Screen Recording (to see the screen). It isn't available in -p mode (non-interactive). The tool names inside the server change from version to version, so the names in the examples below are illustrative: Claude calls them, not you.

2. Through the Anthropic API: you implement the actions yourself in your own environment (Docker, a virtual machine, Xvfb). The current version of the tool, computer_toolset_20260801, doesn't need a beta header; for older models the beta computer_20251124 remains, with the header anthropic-beta: computer-use-2025-11-24.

This lesson focuses on Claude Code, but understanding the API level matters if you want to build your own automations. If Claude Code has a more precise route, Claude picks it on its own: first an MCP server, then Bash, then Claude in Chrome for the web, and screen control only as a last resort.

Three access levels (Tiered Access)

🎨 Picture this: the three access levels work like the badge system in an office building. The lobby (Read): you can look but not touch. The work floors (Click): you can press buttons but can't handle the documents. The server room (Full): full access, for trusted people only.

Computer Use works on the principle of explicit permission. When Claude asks for access to an app (in the tool this is called request_access), you see a dialog that shows the level. The permission lasts for one session (Allow for this session). You can stop Claude at any moment with Esc or Ctrl+C in the terminal. Only one session can control the computer at a time:

Level Applies to What it can do What's blocked
View only (view-only) Browsers (Safari, Chrome, Firefox, Arc) and trading platforms Sees the screen, takes screenshots Clicks and typing are blocked
Clicks only (click-only) Terminals, IDEs (Terminal, iTerm, VS Code, JetBrains) Sees + left click + scroll Typing, right click, drag-and-drop, modifier keys
Full (full) All other apps No restrictions —

The level is set by the app's category, not by your choice. Browsers are always view only, terminals are always clicks only. This is a deliberate security design. On top of that, before you grant permission you see warnings for apps with broad access: terminals and IDEs (equivalent to command-line access), Finder (can read and write any file), System Settings (can change system settings).

Why are browsers view only? For the web there's a more precise and safer tool: the Claude in Chrome extension (it works with the DOM and is faster than pixel clicks; see the lesson Browser automation). Using Computer Use for the web is like hammering nails with a microscope.

Why are terminals on Click? For commands there's the Bash tool, with direct access to the shell. Typing into a terminal through Computer Use would create a risk of commands running out of control.

How to request access

python
# In Claude Code, through an MCP tool
mcp__computer_use__request_access(
    apps=["Finder", "Calendar", "Notes"],
    reason="Automating event creation in Calendar"
)

# The user sees a dialog:
# "Claude wants to control: Finder (full), Calendar (full), Notes (full)"
# "Reason: Automating event creation in Calendar"
# [Allow] [Deny]

Without a request_access call, Claude can't interact with any app. If a new app is needed partway through, another request_access is required.

The working cycle: Screenshot → Analyze → Act (Agentic Loop)

🎨 Picture this: the agentic loop is like someone reading Braille. Touch (screenshot), understand where you are (analysis), press the right spot (action), check the result (next screenshot). Step by step, with confidence.

Every Computer Use action goes through a cycle that Anthropic's documentation calls the agentic loop:

Code
1. screenshot() → get an image of the screen
2. Claude analyzes: what does it see? where is the element it needs? what are the coordinates?
3. Claude requests an action: click, type, scroll
4. The system performs the action
5. screenshot() → Claude checks the result
6. Repeat until the task is done

Coordinates: Claude works with pixel coordinates [x, y] on the screenshot. The coordinate [0, 0] is the top-left corner of the screen.

An example in Claude Code:

python
# Step 1: look at the screen
screen = mcp__computer_use__screenshot()

# Step 2: Claude analyzes the image and finds the element
# Claude sees: the "New Event" button in Calendar, coordinates [1180, 45]

# Step 3: click at the coordinates it found
mcp__computer_use__left_click(coordinate=[1180, 45])

# Step 4: type text
mcp__computer_use__type(text="Client meeting")

# Step 5: screenshot to check the result
screen_after = mcp__computer_use__screenshot()

Available actions

Basic (all versions):

  • screenshot: capture the current screen
  • left_click: click at coordinates [x, y]
  • type: type text
  • key: press a key or key combination (for example "ctrl+s", "Return", "Tab")
  • mouse_move: move the cursor

Extended (in later versions of the tool):

  • scroll: scroll in any direction, with control over the amount
  • left_click_drag: drag and drop between coordinates
  • right_click, middle_click: additional mouse buttons
  • double_click, triple_click: double / triple click
  • left_mouse_down, left_mouse_up: press and release the button separately
  • hold_key: hold a key down for a set time (in seconds)
  • wait: pause between actions

Screen zoom (zoom):

  • zoom: enlarges part of the screen for a close look at small text / UI elements. In the current toolset computer_toolset_20260801 it's on by default; in the earlier beta computer_20251124 (for older models; the official documentation lists which ones) it requires enable_zoom: true in the tool definition

Click modifiers: you can hold Shift / Ctrl / Alt / Cmd (Super) while clicking:

json
{
  "action": "left_click",
  "coordinate": [500, 300],
  "text": "shift"
}

Batched actions: computer_batch

🎨 Picture this: computer_batch is like handing a courier a list of 6 addresses at once instead of calling after every delivery. One trip, six errands. Without batch, it's six round trips.

Instead of six separate round-trip calls, one batch:

python
mcp__computer_use__computer_batch(actions=[
    {"action": "screenshot"},
    {"action": "left_click", "coordinate": [1180, 45]},
    {"action": "type", "text": "Client meeting"},
    {"action": "key", "text": "Tab"},
    {"action": "type", "text": "2:00 PM"},
    {"action": "screenshot"}
])

This is much faster: one round trip instead of six. The actions run in order and stop at the first error. Coordinates inside a batch refer to the screenshot taken before the batch starts.

Teach Mode: interactive teaching

🎨 Picture this: teach mode is like a driving instructor in a dual-control car. They point at the wheel ("turn here"), you press Next, they turn, and you see the result. You learn by watching real actions, not by reading a manual.

Teach mode is a mode in which Claude shows step-by-step hints (tooltips) on top of the screen. The user presses "Next" at each step, and Claude explains what's happening and performs the action.

python
# Request teaching mode
mcp__computer_use__request_teach_access(
    apps=["System Settings"],
    reason="Wi-Fi setup"
)

# Initial screenshot
mcp__computer_use__screenshot()

# Each step: a hint + an action
mcp__computer_use__teach_step(
    explanation="This is the System Settings icon. Let's click it to open settings.",
    next_preview="The System Settings window will open",
    anchor=[512, 400],
    actions=[
        {"action": "left_click", "coordinate": [512, 400]}
    ]
)

In teach mode:

  • Claude's main window hides, and a full-screen overlay with hints appears
  • The user sees an arrow pointing to the element + explanation text
  • Pressing "Next" performs the action and shows the result
  • Pressing "Exit" ends the lesson
  • teach_batch: a batch of steps in one call (faster than one at a time)

It's useful when you're recording a demo video, training your team or showing a client how to set something up.

Zoom: enlarging part of the screen

For inspecting small text, button labels and UI details that are hard to see in a screenshot:

python
# Enlarge the region [x1, y1, x2, y2]
mcp__computer_use__zoom(region=[100, 200, 400, 350])

Zoom returns a full-resolution image of the given region. Coordinates in later clicks always refer to the full-screen screenshot, not the enlarged image.

Practical use cases

Automating native apps that have no API:

python
# Creating a task in Apple Reminders
# (Reminders has no public API for Claude Code)
mcp__computer_use__request_access(
    apps=["Reminders"],
    reason="Creating a task"
)
mcp__computer_use__open_application(app="Reminders")
screen = mcp__computer_use__screenshot()
# Claude finds the "+" button and clicks it
mcp__computer_use__left_click(coordinate=[PLUS_X, PLUS_Y])
mcp__computer_use__type(text="Call the client")
mcp__computer_use__key(text="Return")

Adjusting System Settings:

python
# Changing options in macOS System Settings
mcp__computer_use__request_access(
    apps=["System Settings"],
    reason="Adjusting display settings"
)
mcp__computer_use__open_application(app="System Settings")
# Claude sees the screen, finds the right section, clicks, adjusts

Filling in forms in business systems:

python
# A CRM with no API → fill it in through Computer Use
mcp__computer_use__request_access(
    apps=["CRM App"],
    reason="Filling in a client record"
)
mcp__computer_use__open_application(app="CRM App")
# Claude finds the "Client name" field, clicks it, enters the data
mcp__computer_use__computer_batch(actions=[
    {"action": "left_click", "coordinate": [NAME_X, NAME_Y]},
    {"action": "type", "text": "John Smith"},
    {"action": "key", "text": "Tab"},
    {"action": "type", "text": "john@company.com"},
    {"action": "screenshot"}
])

Computer Use through the Anthropic API (for developers)

If you're building your own app with Computer Use, you need to understand the API level:

python
import anthropic

client = anthropic.Anthropic()

# Current toolset: no beta header needed
response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    tools=[
        {"type": "computer_toolset_20260801"},  # toolset version as of October 2026
        {"type": "bash_20250124", "name": "bash"},
    ],
    messages=[{"role": "user", "content": "Take a screenshot of the desktop"}],
)

# For older models, the beta version remains (see the official docs for which models):
# tools=[{"type": "computer_20251124", "name": "computer",
#         "display_width_px": 1024, "display_height_px": 768}],
# betas=["computer-use-2025-11-24"]  # through client.beta.messages.create

Key differences from Claude Code:

  • You implement the actions yourself (capturing screenshots, clicks, typing)
  • An isolated environment is recommended: a Docker container or a virtual machine
  • Claude returns coordinates relative to the scaled screenshot (at most 2576 px on the long side for current models such as Opus 5.5 and Sonnet 5.5, 1568 px for earlier ones)
  • At high resolutions you need to scale the coordinates back up

Cost: each screenshot adds roughly one to two thousand input tokens (more on high-resolution models), plus a small overhead for the tool description. The documentation advises keeping no more than ~20 images in a request and clearing out screenshot history in batches. Token prices: current prices and versions: What's current.

Limits: what Computer Use doesn't do

Limit Reason
Get past CAPTCHAs and verification A built-in restriction: Claude respects anti-bot systems
Enter passwords and banking details User safety
Carry out financial transactions User safety
Click in a browser (view-only level) Use Claude in Chrome: it works with the DOM and is faster
Type in an IDE/terminal (click-only level) Use the Bash tool directly
Drag-and-drop and right click in click-only A technical restriction of the security level
Create social media accounts An Anthropic restriction against impersonation
Identify people in screenshots A restriction under the Acceptable Use Policy

The exact list of what's prohibited is set by Anthropic's Acceptable Use Policy and the computer use safety guide; check against them.

Accuracy limits (from the official documentation):

  • Latency: Computer Use is slower than a person's normal actions. It's a background tool, not a replacement for your mouse
  • Coordinate accuracy: Claude can get coordinates wrong, especially in complex interfaces
  • Niche apps: reliability drops with uncommon apps or with several apps at once
  • Prompt injection: text on web pages or in images can influence Claude's behavior. Anthropic has built in classifiers to detect injections

🎨 Picture this: TeamViewer, except the operator is an AI. It sees your screen and can click and type, but only in the apps you've explicitly allowed, and with restrictions on sensitive actions. You fully control the list of permissions.

Safety: Anthropic's recommendations

From the official documentation:

  1. Use an isolated environment (a virtual machine / container) with minimal privileges for API integrations. In Claude Code, computer use runs on your real desktop, not in a sandbox: the trust boundary is different from Bash's
  2. Don't give the model access to sensitive data: logins, passwords, banking details
  3. Limit internet access to an allowlist of domains
  4. Review decisions with real-world consequences: financial operations, accepting terms, sending data
  5. Never click links from emails and messages through Computer Use: they may be suspicious
  6. Keep prompt injection in mind: text on the screen may try to give Claude its own command. More on this: defending against prompt injection

🎨 Picture this: Computer Use is a last-mile tool. If there's an API, use the API (faster, more precise). If there's an MCP, use the MCP. Computer Use is for when there's no API, no MCP, but there is a task. A hammer is for nails, not for every job.

Computer Use vs. the alternatives

Task Best tool Why
Web automation Claude in Chrome Works with the DOM, faster, more precise
Working with API services Anthropic API / MCP integrations Direct access, no UI
Native apps with no API Computer Use The only option
Cross-app workflows Computer Use Switches between apps
UI testing Computer Use Sees the real UI the way a user does
Teaching a user Teach Mode Interactive hints
Terminal commands Bash tool Direct access to the shell

Practice

Task: automate creating an event in Calendar

  1. Make sure the computer-use MCP is turned on: it's built into Claude Code but off by default. Type /mcp, select computer-use and choose Enable. You need macOS, a Pro or Max plan and an interactive session; macOS will ask for the Accessibility and Screen Recording permissions

  2. Request access to Calendar:

    Type this into the chat
    Ask Claude: "Open Calendar and create an event"
    Claude will automatically call request_access with the Calendar app
    You'll see a permission dialog: click Allow
  3. Take a first screenshot to make sure Claude can see the screen:

    Type this into the chat
    Claude will call screenshot() and describe what it sees
  4. Create a test event called "Claude check" for tomorrow at 10:00 AM

    • Claude will find the create button, click it and fill in the fields
  5. Watch how Claude uses computer_batch: several actions in one call

  6. Verify with a final screenshot that the event was created

  7. Bonus: Try teach mode. Ask Claude to "show me step by step how to create an event in Calendar." It will switch to teaching mode with hints

Goal: understand how the screenshot → analyze → act cycle works and get a feel for the latency (Computer Use is slower than API integrations, but it works with any app).


Tools and resources


Key takeaways

Computer Use is a feature for controlling native apps that have no API (in Claude Code as of October 2026, a research preview on macOS for Pro and Max). In Claude Code it works through a built-in MCP server; in your own apps, through the API (computer_toolset_20260801).

The three access levels are built-in protection. Browsers are view only (for the web there's Claude in Chrome), terminals are clicks only (for commands there's the Bash tool), everything else gets full access.

Agentic loop: screenshot → analysis → action → screenshot → check. Use computer_batch to cut down on round trips and zoom to read small text.

Teach mode is a unique feature for interactive teaching. Claude shows step-by-step hints on top of the screen, and the user presses Next at each step.

Safety: never give Computer Use access to sensitive data. For API integrations, use Docker/a VM. Claude is trained to resist prompt injection, but extra isolation is a must.


What's next

→ Prompt Caching and the Batch API: economics of scale

The mark stays in this browser only and is never sent anywhere. My progress