Runtime

Live previews for AI-built apps: serve a dev server from a sandbox

Run the app's dev server in the agent's sandbox with spawn, share its port as a preview, and feed the server's errors back to the agent.

Runtime (withruntime.com) gives any port in a sandbox a private HTTPS address with hot reload passing through, and a visit to a paused sandbox's preview was answered in 509 ms at the median, measured from a laptop in the US Mountain time zone. Getting a URL is the easy part. The parts that decide whether an app builder feels alive are the ones around it: a server that survives the agent's restarts, a model that sees the same red error screen the user does, and a preview that sleeps when nobody is looking. This post builds those parts.

What does a preview loop need besides a URL?

Five things, and most products ship with two of them:

Part Without it How to get it
A server that outlives the call The page goes blank when the agent's command returns Start it with spawn, never exec or nohup
Readiness before the link The user's first view is a 502 or a waking page Poll the port inside the sandbox first
Errors the agent can read The agent says "done" while the user stares at a compile error Log to a file; check the page after every edit
A restart when it is needed New packages or config never load; the agent loops on a ghost bug Restart on dependency or config changes only
Sleep between visits You pay for running machines nobody looks at Idle pause; a visit wakes it

The basics of sharing a port, tokens and visibility are in share a port with a preview. This post is about the other four rows.

Which dev servers need a setting to work behind a preview?

Most do, because a preview's address is a name under runtimehost.com and modern dev servers refuse host names they do not know. The refusal protects against DNS rebinding, so allow the one domain rather than every host:

Dev server What to set
Vite, and frameworks on it (SvelteKit, Astro, React Router) __VITE_ADDITIONAL_SERVER_ALLOWED_HOSTS=.runtimehost.com in the server's env (Vite)
Next.js allowedDevOrigins: ["*.runtimehost.com"] in next.config.ts (Next.js)
webpack dev server devServer.allowedHosts: [".runtimehost.com"]
Django runserver ALLOWED_HOSTS = [".runtimehost.com"] in settings
Rails config.hosts << ".runtimehost.com" in config/environments/development.rb
Flask, FastAPI, Express Nothing; they answer any host name

Two settings you do not need. Listening on localhost is enough, because the preview reaches the port from inside the sandbox. And hot reload needs no port or protocol override: the page is served over HTTPS on the standard port, so the dev server's WebSocket follows it there.

Put the setting into your template project once, so the agent never has to discover it from an error page.

How do you start the server so it survives?

Use spawn with the output sent to a log file, and wait for the port before you hand out the link. spawn keeps the process running after your call returns; anything an exec started, nohup ... & included, ends with that command. The log file matters for the next section: it is how the agent will read what the server says.

The tool below is what the agent calls after every batch of edits. It starts the server if nothing is listening, restarts it when asked, waits until the port answers, and returns what the agent needs to judge its own work:

TypeScriptimport { Sandbox } from "withruntime";const APP = "/workspace/app";const PORT = 3000;const LOG = "/workspace/.dev.log";const START = `npm run dev -- --port ${PORT} > ${LOG} 2>&1`;async function listening(sbx: Sandbox) {  const r = await sbx.exec(`curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/`);  return r.stdout.trim() !== "000" && r.stdout.trim() !== "";}// The agent's "check_preview" tool: is the app up, and what is it saying?export async function checkPreview(sbx: Sandbox, opts: { restart?: boolean } = {}) {  if (opts.restart) {    for (const p of await sbx.processes.list())      if (p.state === "running" && p.command.includes("npm run dev"))        await (await sbx.processes.get(p.id)).kill("SIGTERM"); // its whole process group    await sbx.exec(      `for i in $(seq 40); do curl -s -o /dev/null http://localhost:${PORT}/ || exit 0; sleep 0.25; done`,    );  }  if (opts.restart || !(await listening(sbx))) {    await sbx.spawn(START, {      cwd: APP,      env: { __VITE_ADDITIONAL_SERVER_ALLOWED_HOSTS: ".runtimehost.com" },    });    await sbx.exec(      `for i in $(seq 120); do curl -s -o /dev/null http://localhost:${PORT}/ && exit 0; sleep 0.5; done; exit 1`,      { timeoutMs: 90_000 },    );  }  const page = await sbx.exec(`curl -s -w '\\n%{http_code}' http://localhost:${PORT}/`);  const status = page.stdout.trim().split("\n").pop();  const log = await sbx.exec(    `tail -n 40 ${LOG} 2>/dev/null | grep -iE 'error|failed|cannot' || true`,  );  return { status, errors: log.stdout.trim() || "none in the last 40 lines" };}await using sbx = await Sandbox.create({ idlePauseSeconds: 900 });// A stand-in for the agent's project: any app whose `npm run dev` takes --port.await sbx.files.write(`${APP}/package.json`, '{"scripts":{"dev":"node server.js"}}');await sbx.files.write(  `${APP}/server.js`,  'require("http").createServer((q, s) => s.end("ok")).listen(process.argv.at(-1));\n',);console.log(await checkPreview(sbx));
Pythonfrom withruntime import SandboxAPP, PORT, LOG = "/workspace/app", 3000, "/workspace/.dev.log"START = f"npm run dev -- --port {PORT} > {LOG} 2>&1"def listening(sbx) -> bool:    code = sbx.exec(f"curl -s -o /dev/null -w '%{{http_code}}' http://localhost:{PORT}/").stdout.strip()    return code not in ("", "000")def check_preview(sbx, restart: bool = False) -> dict:    """The agent's "check_preview" tool: is the app up, and what is it saying?"""    if restart:        for p in sbx.processes():            if p["state"] == "running" and "npm run dev" in p["command"]:                sbx.process(p["id"]).kill("SIGTERM")  # its whole process group        sbx.exec(f"for i in $(seq 40); do curl -s -o /dev/null http://localhost:{PORT}/ || exit 0; "                 "sleep 0.25; done")    if restart or not listening(sbx):        sbx.spawn(START, cwd=APP, env={"__VITE_ADDITIONAL_SERVER_ALLOWED_HOSTS": ".runtimehost.com"})        sbx.exec(f"for i in $(seq 120); do curl -s -o /dev/null http://localhost:{PORT}/ && exit 0; "                 "sleep 0.5; done; exit 1", timeout_ms=90_000)    page = sbx.exec(f"curl -s -w '\\n%{{http_code}}' http://localhost:{PORT}/")    status = page.stdout.strip().split("\n")[-1]    log = sbx.exec(f"tail -n 40 {LOG} 2>/dev/null | grep -iE 'error|failed|cannot' || true")    return {"status": status, "errors": log.stdout.strip() or "none in the last 40 lines"}with Sandbox.create(idle_pause_seconds=900) as sbx:    # A stand-in for the agent's project: any app whose `npm run dev` takes --port.    sbx.files.write(f"{APP}/package.json", '{"scripts":{"dev":"node server.js"}}')    sbx.files.write(f"{APP}/server.js",                    'require("http").createServer((q, s) => s.end("ok")).listen(process.argv.at(-1));\n')    print(check_preview(sbx))

The readiness loop runs inside the sandbox, so each poll costs a local connection rather than a round trip from your server.

How does the agent see what the user sees?

By reading the same two things the user's browser shows: the HTTP status and the server's own error output. Most dev servers answer 200 with an error overlay when a module fails to compile, so the status alone lies. The log does not: Vite, Next.js and webpack print the failing file and line there. Returning both to the model after every edit catches the commonest failure of app builders, an agent that reports success while the page is broken.

Two refinements pay for themselves quickly:

  • Return only new log lines. Record the log's size before the edit with stat -c %s, then return what was appended after it, so an error fixed three turns ago does not send the agent back to fix it again.
  • Check the route the agent changed, not just /. Pass the path into the tool. A broken settings page behind a working home page is invisible to a check of the home page.

For errors that only happen in the browser, such as a React render error or a failed fetch, give the agent a headless browser in the same sandbox and return the console's errors; Playwright reaches localhost there (Playwright).

When does the dev server need a restart?

Only when hot reload cannot apply the change, which is a short list: the lockfile or package.json changed, a framework config such as vite.config.ts or next.config.ts changed, or an environment file such as .env changed. Everything else, every component and every style, reloads in place, and restarting for it only throws away the browser's state and costs the user a blank screen.

Make the rule mechanical rather than leaving it to the model. Diff the list of files the agent wrote this turn against those names and pass restart: true when one matches. Agents that may restart freely tend to do it after every edit, and the user watches the preview flicker.

What should the agent's instructions say about the preview?

Exactly what the tool already enforces, so the model does not fight it. A model that was never told the server is managed will try to start its own on another port, and the user's preview keeps showing the old one. Four lines in the system prompt prevent most of that:

textThe app's dev server is already running and the user sees it live. Never start, stopor move a server yourself, and never change the port. After each batch of edits,call check_preview, with restart=true only if you changed dependencies, config or.env. Do not say you are done until check_preview returns status 200 and no errors.

The last sentence does the most work. It turns "done" from the model's opinion into a condition your code can check.

How do you show the preview inside your own product?

Put urlWithToken in an iframe and name your site as the only one allowed to embed it:

TypeScriptimport { Sandbox } from "withruntime";const sbx = await Sandbox.connect(process.env.PROJECT_SANDBOX_ID ?? "");const preview = await sbx.previews.create(3000, {  embedOrigins: ["https://app.example.com"], // no other site may frame it  ttlSeconds: 8 * 3600, // one working day});console.log(`<iframe src="${preview.urlWithToken}"></iframe>`);

The browser keeps the token for that frame on your site alone, so the app's scripts, styles and live reload load without it, and any other site's iframe is refused (previews in the SDK). Because the preview is on runtimehost.com and never on your domain, the agent's app can never read your product's cookies or call your API as the user.

The same separation has one cost to plan for. To the browser, the app inside the frame is a third-party site, and browsers increasingly refuse cookies set inside a cross-site frame. If the agent builds a sign-in flow, it may work in a new tab and fail in the frame. Give users an "open in a new tab" button next to every preview, and tell the agent so in its instructions.

What does a preview cost while nobody looks?

Only storage. With idlePauseSeconds set, the sandbox pauses once requests, commands and traffic stop, keeping its files, memory and the running dev server. Preview requests count as use, so an open tab with live reload keeps it awake, and a closed one lets it sleep. The next visit wakes it behind a short "Waking up" page (wake on request).

While it runs, a 2 vCPU, 4 GiB sandbox costs $0.03125 an hour waiting on edits, because CPU is billed on use at $0.025 per vCPU-hour. Paused, it costs $0.08 per GB of saved state a month (pricing). A worked monthly bill for 200 projects is in preview apps an agent builds.

In short

  • A URL is one part of five: the server must survive, be ready, report errors, restart only when needed and sleep when idle.
  • Allow .runtimehost.com in the dev server's host check once, in the template, rather than per project.
  • After every edit, give the model the page's status and the new lines of the server log; the status alone hides compile errors.
  • Embed with urlWithToken and embedOrigins, and keep an "open in a new tab" button for apps that set cookies.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). On Runtime, each project's dev server runs in its own Firecracker microVM, its preview is private until you share it, and a paused project answers a visit in 509 ms measured from a laptop. Create an account at withruntime.com and spawn your template's dev server in the first sandbox.

400 sandbox hours,every month.

Hours of a 1 GB sandbox, included free.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card