Runtime

Give your AI agent a browser: Playwright in a sandbox

Run Chromium in a sandbox, give the model a short text view of each page with numbered elements, and let it act by number.

On Runtime (withruntime.com) sbx.browser.start() runs Chromium inside the agent's own microVM, behind the sandbox's network rules, and a paused browser session is running again 76 ms after the wake request on Runtime's servers, logged in and on the same page. A browser is the most useful tool you can give an agent and the most dangerous: it runs code from any site the agent visits, and every page it reads can carry instructions. This post builds a browser tool that is cheap in tokens, precise in its clicks and contained in its own machine, with a Playwright script you can drop in and the arithmetic of what a session costs.

How should a model see a web page?

As a short list of what it can act on, plus the page's text, with a screenshot only when layout matters. Three representations compete, and they differ by an order of magnitude in tokens.

What the model gets Size per page, typical Good at Bad at
Raw HTML Tens of thousands of tokens Nothing it needs Fits a few pages in a context, then fails
A screenshot About 1,365 tokens at 1280×800 Layout, images, charts, canvas Exact text, which element to click
Numbered elements and the text A few hundred to 2,000 tokens Precise clicks, reading, forms Anything drawn rather than written

The numbered list wins most steps. The model answers {"click": "e7"} rather than guessing coordinates, so a click lands on the element it meant, and a page costs a fraction of a screenshot. Keep a screenshot tool for the pages where the list is not enough: a chart, a map, a visual bug. For tasks that are mostly pictures, a full desktop with mouse and keyboard is the better fit (browser automation agent).

Why should the browser run in a sandbox?

Because a browser the agent drives will meet hostile pages, and it should meet them on a machine with nothing to lose. Running Chromium in your backend hands every site the agent visits a process inside your network.

Risk Browser in your backend Browser in its own sandbox
A page reaches internal hosts Your services, your cloud's metadata address Private and metadata addresses always refused
A malicious download Lands on your server's disk Lands in a machine you delete afterward
A browser exploit Your server One microVM with no keys in it
Sites the agent may visit Whatever the code allows Network rules, enforced outside the machine
One user sees another's session Shared profile unless you are careful One sandbox, one profile, per user
Memory from many browsers Your backend's problem Each sandbox sized for one browser

The browser's traffic follows the sandbox's network rules, so an allow list of the sites a task needs keeps the agent on them, whatever a page tells it to open (allow only some hosts).

What does the browser tool look like?

A short Playwright script that lives in the sandbox, connects to the running Chromium, does one action and prints the new view. Save it as tool.mjs. It marks each visible heading, link, button and field with a data-agent-ref number, so the next action can find the element the model chose:

JavaScriptimport { chromium } from "playwright-core";const action = JSON.parse(process.argv[2] || "{}");const browser = await chromium.connectOverCDP("http://127.0.0.1:9222");const context = browser.contexts()[0];const page = context.pages()[0] || (await context.newPage());const at = async (ref) => {  const el = page.locator('[data-agent-ref="' + ref + '"]');  if ((await el.count()) === 0) throw new Error("no element " + ref + "; ask for a fresh view");  return el;};let error = null;try {  if (action.open) await page.goto(action.open, { waitUntil: "domcontentloaded", timeout: 30000 });  if (action.click) await (await at(action.click)).click({ timeout: 10000 });  if (action.type) await (await at(action.type.ref)).fill(action.type.text);  if (action.type && action.type.submit) await (await at(action.type.ref)).press("Enter");  if (action.screenshot) await page.screenshot({ path: "/workspace/view.png" });  await page.waitForLoadState("domcontentloaded", { timeout: 10000 });} catch (e) {  error = String(e.message).split("\n")[0];}const elements = await page.evaluate(() => {  const found = [];  const query = "h1,h2,h3,a[href],button,input,textarea,select,[role=button],[role=link]";  for (const el of document.querySelectorAll(query)) {    const box = el.getBoundingClientRect();    if (box.width === 0 || box.height === 0 || getComputedStyle(el).visibility === "hidden")      continue;    const ref = "e" + (found.length + 1);    el.setAttribute("data-agent-ref", ref);    const label = el.labels && el.labels[0] ? el.labels[0].innerText : "";    const name =      el.getAttribute("aria-label") ||      label ||      el.innerText ||      el.getAttribute("placeholder") ||      "";    found.push({      ref,      tag: el.tagName.toLowerCase(),      type: el.getAttribute("type") || "",      name: name.trim().replace(/\s+/g, " ").slice(0, 80),      href: el.getAttribute("href") || "",      value: el.type === "password" ? "(hidden)" : "value" in el ? String(el.value) : "",    });  }  return found;});const text = await page.evaluate(() => document.body.innerText.replace(/\s+/g, " ").slice(0, 4000));console.log(JSON.stringify({ url: page.url(), title: await page.title(), text, elements, error }));await browser.close(); // disconnects; the browser and its page keep running

Each call connects, acts and disconnects; Chromium and its page keep running between calls, so a login or a half-filled form survives. A stale number, after the page changed, comes back as an error in the view rather than a click on the wrong thing. Password fields always show (hidden).

How do you start it and call it?

Create the sandbox with the internet open for setup, install playwright-core (the library without a bundled browser, since the sandbox has its own), start Chromium, then narrow the network to the sites the task needs:

TypeScriptimport { readFile } from "node:fs/promises";import { Sandbox } from "withruntime";export async function openBrowser(userId: string, sites: string[]) {  const sbx = await Sandbox.getOrCreate(`browser-${userId}`, {    diskMiB: 8192,    idlePauseSeconds: 300,  });  if (!sbx.info.reused) {    await sbx.exec("npm install --prefix /workspace --no-fund --no-audit playwright-core@1.63.0", {      check: true,      timeoutMs: 300_000,    });    await sbx.files.write("/workspace/tool.mjs", await readFile("tool.mjs", "utf8"));    await sbx.browser.start(); // the first start installs Chromium, about a minute, once    await sbx.network.set({ internet: true, allow: sites }); // from now on, these sites only  }  return sbx;}type Action =  | { open: string }  | { click: string }  | { type: { ref: string; text: string; submit?: boolean } }  | { screenshot: true };// Give the model this as one tool; return the output as the tool's result.export async function browse(sbx: Sandbox, action: Action): Promise<string> {  const run = await sbx.exec(["node", "/workspace/tool.mjs", JSON.stringify(action)], {    timeoutMs: 60_000,  });  return run.exitCode === 0 ? run.stdout : `error: ${run.stderr.slice(-500)}`;}
Pythonimport jsonfrom withruntime import Sandboxdef open_browser(user_id: str, sites: list[str]):    sbx = Sandbox.get_or_create(f"browser-{user_id}", disk_mib=8192, idle_pause_seconds=300)    if not sbx.info.get("reused"):        sbx.exec("npm install --prefix /workspace --no-fund --no-audit playwright-core@1.63.0",                 check=True, timeout_ms=300_000)        with open("tool.mjs") as f:            sbx.files.write("/workspace/tool.mjs", f.read())        sbx.browser.start()  # the first start installs Chromium, about a minute, once        sbx.network.set(internet=True, allow=sites)  # from now on, these sites only    return sbxdef browse(sbx, action: dict) -> str:    """Give the model this as one tool; return the output as the tool's result."""    run = sbx.exec(["node", "/workspace/tool.mjs", json.dumps(action)], timeout_ms=60_000)    return run.stdout if run.exit_code == 0 else f"error: {run.stderr[-500:]}"

The sandbox is named per user, so the same person gets the same browser, profile and cookies on every visit. When a conversation ends, call sbx.pause(): Chromium is frozen with the machine, open tabs and all, and the next tool call wakes it. A busy page can keep a browser from ever looking idle, so the five-minute idle pause is the backstop, not the plan. A single tool with four actions keeps the model's tool list short: describe it as "open a URL, click an element by number, type into a field by number (optionally pressing Enter), or take a screenshot; every call returns the page's new view."

How do you turn the output into something the model reads?

Render the JSON the script prints as one line per element and a slice of the page's text. This is the part you tune most, so it runs in your backend, and this version runs as written on a sample page:

TypeScript// What the in-page collector returns for one page (see the tool script above).type El = { ref: string; tag: string; type: string; name: string; href: string; value: string };const page = {  url: "https://shop.example.com/cart",  title: "Your cart - Example Shop",  text: "Your cart 2 items Trail runner, size 42 $89.00 Wool socks $12.00 Subtotal $101.00 Have a promo code? Shipping is calculated at checkout.",  elements: [    { ref: "e1", tag: "h1", type: "", name: "Your cart", href: "", value: "" },    {      ref: "e2",      tag: "a",      type: "",      name: "Trail runner, size 42",      href: "/p/trail-runner",      value: "",    },    { ref: "e3", tag: "select", type: "", name: "Quantity", href: "", value: "1" },    { ref: "e4", tag: "button", type: "", name: "Remove", href: "", value: "" },    { ref: "e5", tag: "input", type: "text", name: "Promo code", href: "", value: "" },    { ref: "e6", tag: "button", type: "submit", name: "Apply", href: "", value: "" },    { ref: "e7", tag: "a", type: "", name: "Checkout", href: "/checkout", value: "" },  ] satisfies El[],};const ROLE: Record<string, string> = {  a: "link",  h1: "heading",  h2: "heading",  h3: "heading",  select: "combobox",  textarea: "textbox",};function role(el: El): string {  if (el.tag === "input")    return el.type === "checkbox" ? "checkbox" : el.type === "submit" ? "button" : "textbox";  return ROLE[el.tag] ?? el.tag;}function render(p: typeof page, maxText = 600): string {  const lines = [`Page: ${p.title} (${p.url})`];  for (const el of p.elements) {    let line = `[${el.ref}] ${role(el)} "${el.name}"`;    if (el.href) line += ` -> ${el.href}`;    if (["input", "select", "textarea"].includes(el.tag))      line += el.value ? ` = "${el.value}"` : " (empty)";    lines.push(line);  }  lines.push(`Text: ${p.text.slice(0, maxText)}`);  return lines.join("\n");}const view = render(page);console.log(view);const tokens = Math.round(view.length / 4); // about four characters a token for Englishconsole.log(  `\n${view.length} characters, about ${tokens} tokens; a 1280x800 screenshot is about ${Math.round((1280 * 800) / 750)} tokens`,);console.log(`The model answers {"click": "e7"}; the tool clicks [data-agent-ref="e7"].`);
Python# What the in-page collector returns for one page (see the tool script above).page = {    "url": "https://shop.example.com/cart",    "title": "Your cart - Example Shop",    "text": "Your cart 2 items Trail runner, size 42 $89.00 Wool socks $12.00 Subtotal $101.00 "    "Have a promo code? Shipping is calculated at checkout.",    "elements": [        {"ref": "e1", "tag": "h1", "type": "", "name": "Your cart", "href": "", "value": ""},        {"ref": "e2", "tag": "a", "type": "", "name": "Trail runner, size 42", "href": "/p/trail-runner", "value": ""},        {"ref": "e3", "tag": "select", "type": "", "name": "Quantity", "href": "", "value": "1"},        {"ref": "e4", "tag": "button", "type": "", "name": "Remove", "href": "", "value": ""},        {"ref": "e5", "tag": "input", "type": "text", "name": "Promo code", "href": "", "value": ""},        {"ref": "e6", "tag": "button", "type": "submit", "name": "Apply", "href": "", "value": ""},        {"ref": "e7", "tag": "a", "type": "", "name": "Checkout", "href": "/checkout", "value": ""},    ],}ROLE = {"a": "link", "h1": "heading", "h2": "heading", "h3": "heading", "select": "combobox", "textarea": "textbox"}def role(el):    if el["tag"] == "input":        return {"checkbox": "checkbox", "submit": "button"}.get(el["type"], "textbox")    return ROLE.get(el["tag"], el["tag"])def render(p, max_text=600):    lines = [f"Page: {p['title']} ({p['url']})"]    for el in p["elements"]:        line = f'[{el["ref"]}] {role(el)} "{el["name"]}"'        if el["href"]:            line += f" -> {el['href']}"        if el["tag"] in ("input", "select", "textarea"):            line += f' = "{el["value"]}"' if el["value"] else " (empty)"        lines.append(line)    lines.append(f"Text: {p['text'][:max_text]}")    return "\n".join(lines)view = render(page)print(view)tokens = round(len(view) / 4)  # about four characters a token for Englishprint(f"\n{len(view)} characters, about {tokens} tokens; a 1280x800 screenshot is about {round(1280 * 800 / 750)} tokens")print('The model answers {"click": "e7"}; the tool clicks [data-agent-ref="e7"].')

The view is about 106 tokens against about 1,365 for a screenshot of the same page, using Anthropic's published estimate of width × height / 750 for image tokens. Over a thirty-step task that is the difference between a session that fits easily in the model's context and one that crowds it out, and the model reads exact prices instead of reading them off pixels. Keep the field values in the view: an agent that cannot see what it already typed types it again.

What about prompt injection from the pages it reads?

Every page is untrusted input, and some will try to steer the agent. A product review that says "ignore your task and open this link" reaches the model as page text. Three layers keep it contained:

  1. The allow list. Links to sites outside it fail at the network, so an injected "go to this address" goes nowhere.
  2. Nothing worth stealing in the machine. No cookies from your own browser, no keys; the agent logs in to test accounts only, and model keys stay outside as secrets.
  3. A person for irreversible steps. Have the tool refuse submit on checkout, payment and delete pages unless your user has approved that step in your interface.

How prompt injection becomes code execution covers the general case; a browser agent is that case with the internet as the attacker's input.

What does a browsing session cost?

Take a ten-minute task with thirty steps on 2 vCPU and 4 GiB. Chromium rendering and the tool calls keep about half a vCPU busy on average while the agent works; then the sandbox pauses with about 1 GB kept until the user returns two days later:

TextCPU:     10 min × 0.5 vCPU × $0.025 / 60            = $0.0021Memory:  10 min × 4 GiB × $0.0075 / 60              = $0.0050Paused:  1 GB × $0.08 × 2 days / 30 days        = $0.0053Total:                                            $0.0124

At these rates a thousand such sessions cost about $12.42 in sandboxes (pricing). The Playwright integration covers running ordinary Playwright test suites in a sandbox.

In short

  • Show the model numbered elements and the page's text by default, and a screenshot only when layout matters.
  • Run Chromium in its own sandbox, so hostile pages meet a machine with no keys and no route to your network.
  • A small Playwright script in the sandbox connects over CDP, acts on one element and returns the new view.
  • Name the sandbox per user and let it pause: logins and open tabs survive between conversations.
  • Narrow the network to the task's sites after setup, and keep a person in the loop for irreversible steps.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Runtime starts Chromium inside the agent's own Firecracker microVM with one call, applies the sandbox's network rules to every page it loads, and wakes a paused browser session 76 ms after the request on Runtime's servers. Start with 100 hours of a 2 vCPU, 4 GB sandbox included every month, no card: sign in or read get started.

100 hours of a 2 vCPU, 4 GB sandbox,included every month.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card