# Computer use on a Linux desktop: run a computer-use agent safely Start a desktop in a sandbox, send its screenshots to the model, and turn each action it returns into one click, key press or screenshot. **Runtime (withruntime.com) runs a full Linux desktop inside each sandbox's Firecracker microVM, and a sandbox started from a snapshot of a ready desktop is running 440 ms after the request on Runtime's servers.** Some software has no API: an old internal admin tool, a desktop application, a vendor portal built for people with a mouse. A computer-use agent works it the way a person would, from screenshots, and that makes it the most powerful kind of agent and the one most in need of a machine of its own. This post connects Claude's computer toolset to a sandbox desktop, handles the parts that trip up first versions, and sets the limits that make it safe to leave running. ## When does an agent need a desktop, not a browser API? Less often than it seems. A screenshot loop is slower and costs more tokens than any structured interface, so use it where nothing else reaches: | Interface | Speed per step | What it reaches | Use it for | | ----------------------------- | -------------- | --------------------------------------- | -------------------------------------- | | An API or a CLI | Milliseconds | What the vendor exposes | Everything it covers | | A browser over CDP | Under a second | Web pages, through the page's structure | Web apps, forms, scraping | | A desktop through screenshots | Seconds | Anything drawn on a screen | Desktop apps, canvas apps, odd portals | Many teams end up with both lower rows: a CDP browser for the web steps and a desktop for the one application that has only a window. They can share a sandbox, since a browser started with `headless: false` appears on the desktop. ## How does the computer toolset map onto a sandbox desktop? Claude's computer toolset, `computer_toolset_20260801`, gives the model member tools such as `screenshot`, `left_click` and `type`. The model calls them as ordinary tool calls, often several in one reply, and your code performs each one. Almost every member is one call on Runtime's desktop: | Member tool | Runtime call | Notes | | ------------------------------------------- | ------------------------------------------ | ---------------------------------------- | | `screenshot` | `desktop.screenshot()` | PNG bytes, returned as an image | | `left_click`, `right_click`, `middle_click` | `desktop.click(x, y, { button })` | At `coordinate`, or where the pointer is | | `double_click` | `desktop.doubleClick(x, y)` | | | `mouse_move` | `desktop.move(x, y)` | | | `left_click_drag` | `desktop.drag(start, end)` | | | `left_mouse_down`, `left_mouse_up` | `desktop.mouseDown()`, `desktop.mouseUp()` | For drags the model builds itself | | `type` | `desktop.type(text)` | Literal text | | `key` | `desktop.press(keys)` | Same key names: `Return`, `ctrl+s` | | `scroll` | `desktop.scroll(dy, { dx, x, y })` | Direction and amount become wheel clicks | | `cursor_position` | `desktop.cursor()` | Returned as text | | `wait` | A sleep in your code | | | `zoom`, `triple_click`, `hold_key` | Turned off in the request | The toolset's `configs` disables them | Turning a member off is better than half-supporting it. The model then never asks for it, instead of asking and getting an error. ## What does the whole agent look like? The loop is the ordinary tool loop with three details of its own: results carry `toolset_name: "computer"`, a screenshot goes back as an image, and when one action in a batch fails, the rest are reported as not run: ```ts check import { writeFile } from "node:fs/promises"; import Anthropic from "@anthropic-ai/sdk"; import { Sandbox } from "withruntime"; const client = new Anthropic(); // reads ANTHROPIC_API_KEY const SYSTEM = "You operate a Linux desktop. Before any purchase, deletion, message to a person or " + "acceptance of terms, stop and describe what you are about to do instead of doing it."; type Point = [number, number]; type Input = { coordinate?: Point; start_coordinate?: Point; text?: string; repeat?: number; scroll_direction?: "up" | "down" | "left" | "right"; scroll_amount?: number; duration?: number; }; // One member tool call on the sandbox's desktop: a screenshot, a line of text, or "OK". async function act(desktop: Sandbox["desktop"], name: string, input: Input) { const [x, y] = input.coordinate ?? []; if (input.text && name.endsWith("click")) throw new Error("clicks with held keys are not supported"); const n = input.scroll_amount ?? 3; const dir = input.scroll_direction; switch (name) { case "screenshot": return desktop.screenshot(); case "left_click": await desktop.click(x, y); break; case "right_click": await desktop.rightClick(x, y); break; case "middle_click": await desktop.click(x, y, { button: "middle" }); break; case "double_click": await desktop.doubleClick(x, y); break; case "mouse_move": await desktop.move(x!, y!); break; case "left_click_drag": await desktop.drag(input.start_coordinate!, input.coordinate!); break; case "left_mouse_down": await desktop.mouseDown(); break; case "left_mouse_up": await desktop.mouseUp(); break; case "type": await desktop.type(input.text!); break; case "key": for (let i = 0; i < (input.repeat ?? 1); i++) await desktop.press(input.text!); break; case "scroll": await desktop.scroll(dir === "up" ? -n : dir === "down" ? n : 0, { dx: dir === "left" ? -n : dir === "right" ? n : 0, ...(x === undefined ? {} : { x, y }), }); break; case "cursor_position": { const c = await desktop.cursor(); return `X=${c.x}, Y=${c.y}`; } case "wait": await new Promise((r) => setTimeout(r, Math.min(input.duration ?? 1, 30) * 1000)); break; default: throw new Error(`${name} is not available`); } return "OK"; } export async function computerUse(goal: string, maxTurns = 60): Promise { await using sbx = await Sandbox.create({ diskMiB: 8192, timeoutSeconds: 3600, network: { internet: true, allow: ["portal.example.com", "*.example.com"] }, }); const { streamUrl } = await sbx.desktop.start({ width: 1280, height: 800 }); console.log("watch it live:", streamUrl); // private: the link carries its token const recording = await sbx.desktop.recordings.start({ fps: 5, maxMiB: 256 }); await sbx.desktop.open("https://portal.example.com"); const messages: Anthropic.MessageParam[] = [{ role: "user", content: goal }]; try { for (let turn = 0; turn < maxTurns; turn++) { const reply = await client.messages.create({ model: "claude-opus-5-5", max_tokens: 16_000, system: SYSTEM, tools: [ { type: "computer_toolset_20260801", configs: { zoom: { enabled: false }, triple_click: { enabled: false }, hold_key: { enabled: false }, }, }, ], messages, }); messages.push({ role: "assistant", content: reply.content }); if (reply.stop_reason !== "tool_use") return reply.content.flatMap((b) => (b.type === "text" ? [b.text] : [])).join("\n"); const results: Anthropic.ToolResultBlockParam[] = []; let failed = false; for (const call of reply.content) { if (call.type !== "tool_use") continue; const result = { type: "tool_result" as const, tool_use_id: call.id, toolset_name: "computer", }; if (failed) { results.push({ ...result, is_error: true, content: "Not executed: an earlier computer action in this turn failed.", }); continue; } try { const out = await act(sbx.desktop, call.name, call.input as Input); const data = typeof out === "string" ? "" : Buffer.from(out).toString("base64"); results.push({ ...result, content: typeof out === "string" ? [{ type: "text", text: out }] : [{ type: "image", source: { type: "base64", media_type: "image/png", data } }], }); } catch (error) { failed = true; results.push({ ...result, is_error: true, content: `Error: ${(error as Error).message}`, }); } } messages.push({ role: "user", content: results }); } return "Stopped at the turn limit."; } finally { await sbx.desktop.recordings.stop(recording.id); await writeFile( `computer-use-${sbx.id}.mp4`, await sbx.desktop.recordings.download(recording.id), ); } } ``` ```python check import base64 import time import anthropic from withruntime import Sandbox client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY SYSTEM = ("You operate a Linux desktop. Before any purchase, deletion, message to a person or " "acceptance of terms, stop and describe what you are about to do instead of doing it.") TOOLS = [{"type": "computer_toolset_20260801", "configs": {"zoom": {"enabled": False}, "triple_click": {"enabled": False}, "hold_key": {"enabled": False}}}] def act(desktop, name: str, i: dict): """One member tool call on the sandbox's desktop: a screenshot, a line of text, or "OK".""" x, y = i.get("coordinate") or (None, None) if i.get("text") and name.endswith("click"): raise ValueError("clicks with held keys are not supported") n, direction = i.get("scroll_amount", 3), i.get("scroll_direction") if name == "screenshot": return desktop.screenshot() if name == "left_click": desktop.click(x, y) elif name == "right_click": desktop.right_click(x, y) elif name == "middle_click": desktop.click(x, y, button="middle") elif name == "double_click": desktop.double_click(x, y) elif name == "mouse_move": desktop.move(x, y) elif name == "left_click_drag": desktop.drag(tuple(i["start_coordinate"]), tuple(i["coordinate"])) elif name == "left_mouse_down": desktop.mouse_down() elif name == "left_mouse_up": desktop.mouse_up() elif name == "type": desktop.type(i["text"]) elif name == "key": for _ in range(i.get("repeat", 1)): desktop.press(i["text"]) elif name == "scroll": desktop.scroll({"up": -n, "down": n}.get(direction, 0), dx={"left": -n, "right": n}.get(direction, 0), x=x, y=y) elif name == "cursor_position": c = desktop.cursor() return f"X={c['x']}, Y={c['y']}" elif name == "wait": time.sleep(min(i.get("duration", 1), 30)) else: raise ValueError(f"{name} is not available") return "OK" def computer_use(goal: str, max_turns: int = 60) -> str: with Sandbox.create(disk_mib=8192, timeout_seconds=3600, network={"internet": True, "allow": ["portal.example.com", "*.example.com"]}) as sbx: print("watch it live:", sbx.desktop.start(width=1280, height=800)["streamUrl"]) recording = sbx.desktop.recordings.start(fps=5, max_mib=256) sbx.desktop.open("https://portal.example.com") messages: list[dict] = [{"role": "user", "content": goal}] try: for _ in range(max_turns): reply = client.messages.create(model="claude-opus-5-5", max_tokens=16_000, system=SYSTEM, tools=TOOLS, messages=messages) messages.append({"role": "assistant", "content": reply.content}) if reply.stop_reason != "tool_use": return "\n".join(b.text for b in reply.content if b.type == "text") results, failed = [], False for call in (b for b in reply.content if b.type == "tool_use"): result = {"type": "tool_result", "tool_use_id": call.id, "toolset_name": "computer"} if failed: results.append({**result, "is_error": True, "content": "Not executed: an earlier computer action in this turn failed."}) continue try: out = act(sbx.desktop, call.name, call.input) except Exception as error: failed = True results.append({**result, "is_error": True, "content": f"Error: {error}"}) continue if isinstance(out, bytes): image = {"type": "base64", "media_type": "image/png", "data": base64.b64encode(out).decode()} results.append({**result, "content": [{"type": "image", "source": image}]}) else: results.append({**result, "content": [{"type": "text", "text": out}]}) messages.append({"role": "user", "content": results}) return "Stopped at the turn limit." finally: sbx.desktop.recordings.stop(recording["id"]) with open(f"computer-use-{sbx.id}.mp4", "wb") as f: f.write(sbx.desktop.recordings.download(recording["id"])) ``` ## Why is the screen exactly 1280 by 800? Because the model's clicks are in the coordinates of the screenshots you send, and the API does not shrink images for you. A screenshot larger than the model's image limit is refused, and if you shrink it yourself, every coordinate that comes back has to be scaled up again before the click. Miss that once and the agent clicks a little left of every button, which looks like a model failing rather than a bug in your code. The simplest fix is to make the screen the size you send. 1280 by 800 is well inside the limit for current Claude models, matches how most web applications are designed, and keeps each screenshot to a modest number of tokens. Coordinates then pass straight through with no arithmetic. Larger screens make each step slower and more expensive without making the model more accurate. ## What should a batch of actions do when one fails? Stop. The model often sends a short plan in one reply: click the search box, type a query, press Return, take a screenshot. If the click fails, typing the query lands somewhere else, and pressing Return may submit a form nobody meant to submit. The loop above runs the calls in order, and after the first error answers the rest with "Not executed", so every call still gets a result and the model looks again before trying anything else. ## How long can one run go on? Every screenshot stays in the transcript, so a long task grows its context by an image per step. The turn limit is the first answer: a task that needs more than sixty steps is usually two tasks. Do not delete old screenshots from the messages yourself. On current Claude models, editing earlier turns invalidates the model's later reasoning; use the API's own tool result clearing, which removes old results on the server side, when a run must go longer. ## How do you keep a computer-use agent from doing harm? The model acts on whatever it sees, and what it sees includes text written by strangers: a web page, an email in the portal, a document. Instructions in that text can steer it ([prompt injection to code execution](/blog/prompt-injection-to-code-execution)). Limit what a steered agent could reach: - **Its own machine.** The desktop is inside the sandbox's microVM, so a download it opens or a program it starts never touches yours. - **A short network allow list.** Name the sites the task needs and nothing else, so a page cannot send it somewhere new ([allow only some hosts](/how-to/allow-only-some-hosts)). - **No logged-in browser of yours.** Sign in with an account made for the agent, with the least access the task needs. - **A person for the irreversible.** The system prompt tells the model to stop and describe before a purchase, a deletion, a message to a person or new terms. Your code returns that description to a human, not to the next turn. - **A record.** The recording is an MP4 of every step, kept on the sandbox's disk until your code downloads it. When the agent does something odd, you can watch exactly what it saw. The live link lets you watch while it runs. ## How fast is it, and what does it cost? The first desktop start in a sandbox installs it, about 90 seconds and 1 GB of disk, once. Do that once, then take a memory snapshot of the sandbox with the desktop running and the application open, signed in with the agent's account. Each task then starts from the snapshot, already at the first screen, in 440 ms ([snapshots and forks](/docs/javascript#snapshots-and-forks)). After that, the model sets the pace. Each step is one model call with a screenshot in it, usually a few seconds, while the desktop's own actions return in a fraction of that. A 40-step task takes a few minutes, during which the sandbox mostly waits: CPU is billed only when used, at $0.025 per vCPU-hour, and memory at $0.0075 per GiB-hour while it runs. The model's image tokens are the larger cost, which is one more reason to keep the screen at 1280 by 800 rather than larger. ## How do you test it without spending tokens? Replay. Save the model's replies from one good run, then feed them back through `act` against a fresh sandbox from the same snapshot and compare the final screenshot with the original. A replay that diverges means the application changed, which is exactly what you want to learn before tonight's real run. The same idea, applied to whole agent runs, is in [reproducible agent evals](/blog/reproducible-agent-evals). ## In short - Use a desktop only where no API or browser interface reaches; it is the slowest and most capable interface. - Map each computer toolset member to one desktop call, and turn off the members you do not support. - Make the screen the size you send, so coordinates need no scaling. - Stop a batch at its first failed action and answer the rest as not run. - Contain it: its own microVM, a short allow list, an account made for it, a person for irreversible steps, and a recording of every run. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Every Runtime sandbox can run a desktop, record it to MP4 and share a private live view. [Read how to control a desktop](/how-to/control-a-desktop), then [create an account at withruntime.com](/sign-in) and point the agent above at the application that has no API.