# Give your AI agent a browser: Playwright in a sandbox Run Chromium in a sandbox, give the model a short text view of each page with numbered elements, and let it act by number. **On Runtime (withruntime.com) `sbx.browser.start()` runs Chromium inside the agent's own microVM, behind the sandbox's network rules, and a paused browser session is running again 76 ms after the wake request on Runtime's servers, logged in and on the same page.** A browser is the most useful tool you can give an agent and the most dangerous: it runs code from any site the agent visits, and every page it reads can carry instructions. This post builds a browser tool that is cheap in tokens, precise in its clicks and contained in its own machine, with a Playwright script you can drop in and the arithmetic of what a session costs. ## How should a model see a web page? As a short list of what it can act on, plus the page's text, with a screenshot only when layout matters. Three representations compete, and they differ by an order of magnitude in tokens. | What the model gets | Size per page, typical | Good at | Bad at | | ------------------------------ | ------------------------------ | ------------------------------ | ----------------------------------------- | | Raw HTML | Tens of thousands of tokens | Nothing it needs | Fits a few pages in a context, then fails | | A screenshot | About 1,365 tokens at 1280×800 | Layout, images, charts, canvas | Exact text, which element to click | | Numbered elements and the text | A few hundred to 2,000 tokens | Precise clicks, reading, forms | Anything drawn rather than written | The numbered list wins most steps. The model answers `{"click": "e7"}` rather than guessing coordinates, so a click lands on the element it meant, and a page costs a fraction of a screenshot. Keep a screenshot tool for the pages where the list is not enough: a chart, a map, a visual bug. For tasks that are mostly pictures, a full desktop with mouse and keyboard is the better fit ([browser automation agent](/use-cases/browser-automation-agent)). ## Why should the browser run in a sandbox? Because a browser the agent drives will meet hostile pages, and it should meet them on a machine with nothing to lose. Running Chromium in your backend hands every site the agent visits a process inside your network. | Risk | Browser in your backend | Browser in its own sandbox | | ------------------------------- | -------------------------------------------- | --------------------------------------------- | | A page reaches internal hosts | Your services, your cloud's metadata address | Private and metadata addresses always refused | | A malicious download | Lands on your server's disk | Lands in a machine you delete afterward | | A browser exploit | Your server | One microVM with no keys in it | | Sites the agent may visit | Whatever the code allows | Network rules, enforced outside the machine | | One user sees another's session | Shared profile unless you are careful | One sandbox, one profile, per user | | Memory from many browsers | Your backend's problem | Each sandbox sized for one browser | The browser's traffic follows the sandbox's network rules, so an allow list of the sites a task needs keeps the agent on them, whatever a page tells it to open ([allow only some hosts](/how-to/allow-only-some-hosts)). ## What does the browser tool look like? A short Playwright script that lives in the sandbox, connects to the running Chromium, does one action and prints the new view. Save it as `tool.mjs`. It marks each visible heading, link, button and field with a `data-agent-ref` number, so the next action can find the element the model chose: ```js import { chromium } from "playwright-core"; const action = JSON.parse(process.argv[2] || "{}"); const browser = await chromium.connectOverCDP("http://127.0.0.1:9222"); const context = browser.contexts()[0]; const page = context.pages()[0] || (await context.newPage()); const at = async (ref) => { const el = page.locator('[data-agent-ref="' + ref + '"]'); if ((await el.count()) === 0) throw new Error("no element " + ref + "; ask for a fresh view"); return el; }; let error = null; try { if (action.open) await page.goto(action.open, { waitUntil: "domcontentloaded", timeout: 30000 }); if (action.click) await (await at(action.click)).click({ timeout: 10000 }); if (action.type) await (await at(action.type.ref)).fill(action.type.text); if (action.type && action.type.submit) await (await at(action.type.ref)).press("Enter"); if (action.screenshot) await page.screenshot({ path: "/workspace/view.png" }); await page.waitForLoadState("domcontentloaded", { timeout: 10000 }); } catch (e) { error = String(e.message).split("\n")[0]; } const elements = await page.evaluate(() => { const found = []; const query = "h1,h2,h3,a[href],button,input,textarea,select,[role=button],[role=link]"; for (const el of document.querySelectorAll(query)) { const box = el.getBoundingClientRect(); if (box.width === 0 || box.height === 0 || getComputedStyle(el).visibility === "hidden") continue; const ref = "e" + (found.length + 1); el.setAttribute("data-agent-ref", ref); const label = el.labels && el.labels[0] ? el.labels[0].innerText : ""; const name = el.getAttribute("aria-label") || label || el.innerText || el.getAttribute("placeholder") || ""; found.push({ ref, tag: el.tagName.toLowerCase(), type: el.getAttribute("type") || "", name: name.trim().replace(/\s+/g, " ").slice(0, 80), href: el.getAttribute("href") || "", value: el.type === "password" ? "(hidden)" : "value" in el ? String(el.value) : "", }); } return found; }); const text = await page.evaluate(() => document.body.innerText.replace(/\s+/g, " ").slice(0, 4000)); console.log(JSON.stringify({ url: page.url(), title: await page.title(), text, elements, error })); await browser.close(); // disconnects; the browser and its page keep running ``` Each call connects, acts and disconnects; Chromium and its page keep running between calls, so a login or a half-filled form survives. A stale number, after the page changed, comes back as an error in the view rather than a click on the wrong thing. Password fields always show `(hidden)`. ## How do you start it and call it? Create the sandbox with the internet open for setup, install `playwright-core` (the library without a bundled browser, since the sandbox has its own), start Chromium, then narrow the network to the sites the task needs: ```ts check import { readFile } from "node:fs/promises"; import { Sandbox } from "withruntime"; export async function openBrowser(userId: string, sites: string[]) { const sbx = await Sandbox.getOrCreate(`browser-${userId}`, { diskMiB: 8192, idlePauseSeconds: 300, }); if (!sbx.info.reused) { await sbx.exec("npm install --prefix /workspace --no-fund --no-audit playwright-core@1.63.0", { check: true, timeoutMs: 300_000, }); await sbx.files.write("/workspace/tool.mjs", await readFile("tool.mjs", "utf8")); await sbx.browser.start(); // the first start installs Chromium, about a minute, once await sbx.network.set({ internet: true, allow: sites }); // from now on, these sites only } return sbx; } type Action = | { open: string } | { click: string } | { type: { ref: string; text: string; submit?: boolean } } | { screenshot: true }; // Give the model this as one tool; return the output as the tool's result. export async function browse(sbx: Sandbox, action: Action): Promise { const run = await sbx.exec(["node", "/workspace/tool.mjs", JSON.stringify(action)], { timeoutMs: 60_000, }); return run.exitCode === 0 ? run.stdout : `error: ${run.stderr.slice(-500)}`; } ``` ```python check import json from withruntime import Sandbox def open_browser(user_id: str, sites: list[str]): sbx = Sandbox.get_or_create(f"browser-{user_id}", disk_mib=8192, idle_pause_seconds=300) if not sbx.info.get("reused"): sbx.exec("npm install --prefix /workspace --no-fund --no-audit playwright-core@1.63.0", check=True, timeout_ms=300_000) with open("tool.mjs") as f: sbx.files.write("/workspace/tool.mjs", f.read()) sbx.browser.start() # the first start installs Chromium, about a minute, once sbx.network.set(internet=True, allow=sites) # from now on, these sites only return sbx def browse(sbx, action: dict) -> str: """Give the model this as one tool; return the output as the tool's result.""" run = sbx.exec(["node", "/workspace/tool.mjs", json.dumps(action)], timeout_ms=60_000) return run.stdout if run.exit_code == 0 else f"error: {run.stderr[-500:]}" ``` The sandbox is named per user, so the same person gets the same browser, profile and cookies on every visit. When a conversation ends, call `sbx.pause()`: Chromium is frozen with the machine, open tabs and all, and the next tool call wakes it. A busy page can keep a browser from ever looking idle, so the five-minute idle pause is the backstop, not the plan. A single tool with four actions keeps the model's tool list short: describe it as "open a URL, click an element by number, type into a field by number (optionally pressing Enter), or take a screenshot; every call returns the page's new view." ## How do you turn the output into something the model reads? Render the JSON the script prints as one line per element and a slice of the page's text. This is the part you tune most, so it runs in your backend, and this version runs as written on a sample page: ```ts // What the in-page collector returns for one page (see the tool script above). type El = { ref: string; tag: string; type: string; name: string; href: string; value: string }; const page = { url: "https://shop.example.com/cart", title: "Your cart - Example Shop", text: "Your cart 2 items Trail runner, size 42 $89.00 Wool socks $12.00 Subtotal $101.00 Have a promo code? Shipping is calculated at checkout.", elements: [ { ref: "e1", tag: "h1", type: "", name: "Your cart", href: "", value: "" }, { ref: "e2", tag: "a", type: "", name: "Trail runner, size 42", href: "/p/trail-runner", value: "", }, { ref: "e3", tag: "select", type: "", name: "Quantity", href: "", value: "1" }, { ref: "e4", tag: "button", type: "", name: "Remove", href: "", value: "" }, { ref: "e5", tag: "input", type: "text", name: "Promo code", href: "", value: "" }, { ref: "e6", tag: "button", type: "submit", name: "Apply", href: "", value: "" }, { ref: "e7", tag: "a", type: "", name: "Checkout", href: "/checkout", value: "" }, ] satisfies El[], }; const ROLE: Record = { a: "link", h1: "heading", h2: "heading", h3: "heading", select: "combobox", textarea: "textbox", }; function role(el: El): string { if (el.tag === "input") return el.type === "checkbox" ? "checkbox" : el.type === "submit" ? "button" : "textbox"; return ROLE[el.tag] ?? el.tag; } function render(p: typeof page, maxText = 600): string { const lines = [`Page: ${p.title} (${p.url})`]; for (const el of p.elements) { let line = `[${el.ref}] ${role(el)} "${el.name}"`; if (el.href) line += ` -> ${el.href}`; if (["input", "select", "textarea"].includes(el.tag)) line += el.value ? ` = "${el.value}"` : " (empty)"; lines.push(line); } lines.push(`Text: ${p.text.slice(0, maxText)}`); return lines.join("\n"); } const view = render(page); console.log(view); const tokens = Math.round(view.length / 4); // about four characters a token for English console.log( `\n${view.length} characters, about ${tokens} tokens; a 1280x800 screenshot is about ${Math.round((1280 * 800) / 750)} tokens`, ); console.log(`The model answers {"click": "e7"}; the tool clicks [data-agent-ref="e7"].`); ``` ```python # What the in-page collector returns for one page (see the tool script above). page = { "url": "https://shop.example.com/cart", "title": "Your cart - Example Shop", "text": "Your cart 2 items Trail runner, size 42 $89.00 Wool socks $12.00 Subtotal $101.00 " "Have a promo code? Shipping is calculated at checkout.", "elements": [ {"ref": "e1", "tag": "h1", "type": "", "name": "Your cart", "href": "", "value": ""}, {"ref": "e2", "tag": "a", "type": "", "name": "Trail runner, size 42", "href": "/p/trail-runner", "value": ""}, {"ref": "e3", "tag": "select", "type": "", "name": "Quantity", "href": "", "value": "1"}, {"ref": "e4", "tag": "button", "type": "", "name": "Remove", "href": "", "value": ""}, {"ref": "e5", "tag": "input", "type": "text", "name": "Promo code", "href": "", "value": ""}, {"ref": "e6", "tag": "button", "type": "submit", "name": "Apply", "href": "", "value": ""}, {"ref": "e7", "tag": "a", "type": "", "name": "Checkout", "href": "/checkout", "value": ""}, ], } ROLE = {"a": "link", "h1": "heading", "h2": "heading", "h3": "heading", "select": "combobox", "textarea": "textbox"} def role(el): if el["tag"] == "input": return {"checkbox": "checkbox", "submit": "button"}.get(el["type"], "textbox") return ROLE.get(el["tag"], el["tag"]) def render(p, max_text=600): lines = [f"Page: {p['title']} ({p['url']})"] for el in p["elements"]: line = f'[{el["ref"]}] {role(el)} "{el["name"]}"' if el["href"]: line += f" -> {el['href']}" if el["tag"] in ("input", "select", "textarea"): line += f' = "{el["value"]}"' if el["value"] else " (empty)" lines.append(line) lines.append(f"Text: {p['text'][:max_text]}") return "\n".join(lines) view = render(page) print(view) tokens = round(len(view) / 4) # about four characters a token for English print(f"\n{len(view)} characters, about {tokens} tokens; a 1280x800 screenshot is about {round(1280 * 800 / 750)} tokens") print('The model answers {"click": "e7"}; the tool clicks [data-agent-ref="e7"].') ``` The view is about 106 tokens against about 1,365 for a screenshot of the same page, using Anthropic's published estimate of width × height / 750 for image tokens. Over a thirty-step task that is the difference between a session that fits easily in the model's context and one that crowds it out, and the model reads exact prices instead of reading them off pixels. Keep the field values in the view: an agent that cannot see what it already typed types it again. ## What about prompt injection from the pages it reads? Every page is untrusted input, and some will try to steer the agent. A product review that says "ignore your task and open this link" reaches the model as page text. Three layers keep it contained: 1. **The allow list.** Links to sites outside it fail at the network, so an injected "go to this address" goes nowhere. 2. **Nothing worth stealing in the machine.** No cookies from your own browser, no keys; the agent logs in to test accounts only, and model keys stay outside as [secrets](/docs/security#secrets-sandboxes-never-see). 3. **A person for irreversible steps.** Have the tool refuse `submit` on checkout, payment and delete pages unless your user has approved that step in your interface. [How prompt injection becomes code execution](/blog/prompt-injection-to-code-execution) covers the general case; a browser agent is that case with the internet as the attacker's input. ## What does a browsing session cost? Take a ten-minute task with thirty steps on 2 vCPU and 4 GiB. Chromium rendering and the tool calls keep about half a vCPU busy on average while the agent works; then the sandbox pauses with about 1 GB kept until the user returns two days later: ``` CPU: 10 min × 0.5 vCPU × $0.025 / 60 = $0.0021 Memory: 10 min × 4 GiB × $0.0075 / 60 = $0.0050 Paused: 1 GB × $0.08 × 2 days / 30 days = $0.0053 Total: $0.0124 ``` At these rates a thousand such sessions cost about $12.42 in sandboxes ([pricing](/pricing)). The [Playwright integration](/integrations/playwright) covers running ordinary Playwright test suites in a sandbox. ## In short - Show the model numbered elements and the page's text by default, and a screenshot only when layout matters. - Run Chromium in its own sandbox, so hostile pages meet a machine with no keys and no route to your network. - A small Playwright script in the sandbox connects over CDP, acts on one element and returns the new view. - Name the sandbox per user and let it pause: logins and open tabs survive between conversations. - Narrow the network to the task's sites after setup, and keep a person in the loop for irreversible steps. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Runtime starts Chromium inside the agent's own Firecracker microVM with one call, applies the sandbox's network rules to every page it loads, and wakes a paused browser session 76 ms after the request on Runtime's servers. Start with 100 hours of a 2 vCPU, 4 GB sandbox included every month, no card: [sign in](/sign-in) or read [get started](/docs/start).