# Agent sandbox or serverless function: when each one fits Use a serverless function for short, trusted steps you wrote, and a sandbox for code a model wrote that needs a shell, files and time. **Runtime (withruntime.com) gives each conversation its own Firecracker microVM that runs its first command 221 ms after the create request, pauses itself after 60 seconds with nothing happening, and wakes for the next command in 153 ms, measured on Runtime's servers.** Most agent products need both kinds of compute, and the mistakes come from putting a piece in the wrong one. This post splits a typical agent product into its parts, says where each belongs and why, and shows the twenty lines that connect a function to a sandbox safely. ## What is the real difference between a function and a sandbox? A function is an invocation you hand an event to; a sandbox is a machine you hold for as long as the work needs it. Everything else follows from that unit: | Question | Serverless function | Agent sandbox | | ----------------------------------- | ----------------------------------------------- | ------------------------------------------------- | | What you deploy | Your code, reviewed and built ahead of time | Nothing; code arrives while it runs | | Whose code runs | Yours | The model's, and whatever it installs | | State between calls | Not promised; reload it each time | Files, memory and processes stay | | What it can reach by default | Your cloud account, through the function's role | The internet, or nothing, as you set it | | How you talk to it | One request, one response | Many commands, files and processes over a session | | How long one piece of work may take | Capped by the platform, in minutes | As long as it works, while credit lasts | | Who it serves | Every user, on shared warm instances | One user or one task, then gone | The two rows that matter most for agents are "whose code runs" and "what it can reach". A function is built to run trusted code with your credentials. A sandbox is built to run untrusted code with none. ## Why not run model-written code inside the function? Because the function is already holding your keys. A function runs with an identity: an AWS role, a Google service account, the environment variables your platform injects. Code a model wrote and you `exec` inside that function inherits all of it. One prompt injection, and `env` or the metadata endpoint hands an attacker whatever the function could do, which is usually read and write access to your own data ([how injected text becomes a command](/blog/prompt-injection-to-code-execution)). Three more reasons show up in practice: - **Warm instances are reused.** Platforms keep an instance warm and send it the next request, so a file one user's code wrote in `/tmp` can still be there when the next user's code runs. - **The file system is mostly read-only.** `pip install` at run time either fails or lands in a scratch folder that disappears, so the model cannot add the library it needs. - **The run is capped.** A test suite, a build or a data job that runs past the platform's limit is cut off with no way to resume. A sandbox inverts each of these. It is a fresh machine with its own kernel, nothing from any other request in it, `sudo` to install anything, and no credentials unless you put them there ([run untrusted LLM code](/use-cases/run-untrusted-llm-code)). ## Which parts of an agent product belong where? Split the product into four parts and the answer is usually clear: 1. **The front door**: webhooks from GitHub or Slack, your API, the chat endpoint. Short, stateless, your own code, called by many users. A function fits. 2. **The agent loop**: call the model, read the tool calls, run them, repeat. It is your code, but it spends most of its time waiting on the model. Either fits, depending on how long a task runs (next section). 3. **Tool execution**: the shell commands, Python, tests and installs the model asks for. Untrusted, stateful, sometimes long. A sandbox. 4. **The record**: writing results, transcripts and usage to your database. Trusted, short. A function, or the end of the loop. The common mistake is to merge 2 and 3 into one function, which gives the model's code your loop's credentials. The other is to merge 1 and 3 into one sandbox per request, which pays to start a machine for work that needed none. ## Can a function run the agent loop? Yes, if one task fits inside the platform's time limit, and you pay for the waiting. A loop with ten model calls of twenty seconds each spends over three minutes waiting and seconds working. A function billed on wall-clock time charges its full memory for all of that wait, and a durable workflow engine can split it into steps to stay under the cap. For work that runs longer than the cap, or that you want to survive a redeploy, run the loop inside the sandbox instead. The model's API key goes in as a secret the sandbox can use but never read ([secrets](/docs/security#secrets-sandboxes-never-see)), your function starts the task and returns, and a webhook tells you when it is done ([background agents](/use-cases/background-agents)). The sandbox bills its CPU only while the loop actually computes, and the waiting costs $0.03125 an hour for 2 vCPU and 4 GiB. ## How do you call a sandbox from a function safely? Name the sandbox after the conversation and use `getOrCreate`, so a retry finds the machine the first try made. Function platforms retry failed invocations, and a webhook sender retries when you answer slowly; a handler that calls a plain `create` makes a second sandbox on each retry. A name makes the handler safe to run twice, and it makes a conversation's files, installed packages and running processes still there on the next turn: ```ts check import { Sandbox } from "withruntime"; type Turn = { conversationId: string; code: string }; // Your function's handler: one turn of one conversation. export async function handler({ conversationId, code }: Turn) { const sbx = await Sandbox.getOrCreate(`chat-${conversationId}`, { labels: { app: "chat" }, network: { internet: false }, // the model's code reaches nothing idlePauseSeconds: 120, }); await sbx.files.write("/workspace/turn.py", code); const run = await sbx.exec(["python3", "turn.py"], { timeoutMs: 60_000 }); return { ok: run.exitCode === 0, output: (run.stdout + run.stderr).slice(-8000), timedOut: run.timedOut, }; } console.log(await handler({ conversationId: "c-42", code: "print(sum(range(10)))" })); ``` ```python check from withruntime import Sandbox def handler(conversation_id: str, code: str) -> dict: """Your function's handler: one turn of one conversation.""" sbx = Sandbox.get_or_create( f"chat-{conversation_id}", labels={"app": "chat"}, network={"internet": False}, # the model's code reaches nothing idle_pause_seconds=120, ) sbx.files.write("/workspace/turn.py", code) run = sbx.exec(["python3", "turn.py"], timeout_ms=60_000) return { "ok": run.exit_code == 0, "output": (run.stdout + run.stderr)[-8000:], "timed_out": run.timed_out, } print(handler("c-42", "print(sum(range(10)))")) ``` What this gives you, without any state in the function: - **Nothing to store.** The name is the handle, so the function needs no table mapping conversations to sandbox ids ([find a sandbox by name](/how-to/find-a-sandbox-by-name)). - **Nothing to wake.** Between turns the sandbox pauses itself, and the next `exec` wakes it with its files and memory as they were ([wake on request](/how-to/wake-a-sandbox-on-request)). - **Nothing to clean up per request.** The handler never stops the sandbox. A nightly job calls `runtime.sandboxes.stopAll({ labels: { app: "chat" } })`, or stops just the ones whose conversations are over. For a sandbox that must be new on each call, such as one per job rather than per conversation, pass an idempotency key to `create` instead, derived from the event's id, so a retried event gets the same sandbox back rather than a second one ([retry safely](/how-to/retry-safely-with-idempotency-keys)): ```ts import { randomUUID } from "node:crypto"; import { Runtime } from "withruntime"; const runtime = new Runtime(); // A queue or webhook handler: one sandbox per job, however often the event arrives. export async function onJob(event: { id: string; command: string }) { const sbx = await runtime.sandboxes.create( { labels: { app: "jobs", event: event.id }, timeoutSeconds: 1800 }, { idempotencyKey: `job-${event.id}` }, ); try { const run = await sbx.exec(event.command, { timeoutMs: 600_000 }); return { exitCode: run.exitCode, tail: run.stdout.slice(-2000) }; } finally { await sbx.stop(); } } // The event's id comes from your queue; it is the same every time the event is redelivered. const event = { id: `evt_${randomUUID()}`, command: "python3 -c 'import sys; print(sys.version)'" }; console.log(await onJob(event)); ``` ```python from uuid import uuid4 from withruntime import Runtime runtime = Runtime() def on_job(event_id: str, command: str) -> dict: """A queue or webhook handler: one sandbox per job, however often the event arrives.""" sbx = runtime.sandboxes.create(labels={"app": "jobs", "event": event_id}, timeout_seconds=1800, idempotency_key=f"job-{event_id}") try: run = sbx.exec(command, timeout_ms=600_000) return {"exit_code": run.exit_code, "tail": run.stdout[-2000:]} finally: sbx.stop() # The event's id comes from your queue; it is the same every time the event is redelivered. print(on_job(f"evt_{uuid4()}", "python3 -c 'import sys; print(sys.version)'")) ``` ## How fast is the call from a function? Fast enough for a request path. A command in a running sandbox took 53 ms at the median, measured on Runtime's servers, and a paused one ran its next command 153 ms after the call. A brand-new sandbox ran its first command 221 ms after the create request. Add the network between your function's region and Runtime's. From a laptop in the US Mountain time zone, the first command of a new sandbox came back 331 ms after the request ([speed](/docs/speed)). That means you do not need a pool of warm sandboxes, and the first turn of a conversation can create its sandbox inside the user's request ([what a cold start is](/glossary/cold-start)). ## What does each one cost for agent work? Each bills a different thing, so the cheaper one depends on how much of the time is waiting. A function bills the memory it was given for every millisecond of the invocation. A Runtime sandbox bills memory while it runs, CPU only for what its commands use at $0.025 per vCPU-hour, and only storage while paused ([pricing](/pricing)). Take a chat product with 1,000 conversations a day, each with 15 turns, where the model's code runs 2 CPU-seconds a turn and the sandbox stays awake two minutes after each turn before it pauses: ``` Awake: 1,000 × 15 turns × 2 min / 60 × $0.03125 = $15.63 a day CPU: 1,000 × 15 × 2 s / 3,600 × $0.025 = $0.21 a day Total: $15.83 a day ``` Nearly all of that is the awake window, so it is also the lever: drop `idlePauseSeconds` to 30 and the awake line falls to $3.91, at the cost of a wake on turns that arrive later than that ([pause between turns](/blog/pause-agent-sandboxes-between-turns)). A function can run the same code only if it is your own code, so the comparison that matters is where the untrusted part runs, and in a function it runs beside your credentials ([Lambda against a sandbox, priced](/compare/aws-lambda-vs-agent-sandbox)). ## When is a function the better choice? When the code is yours, the work is short and stateless, and it needs your cloud account's identity: - **Glue between services:** parse a webhook, call an API, write a row. - **High fan-out over your own code:** resize ten thousand images with a library you deployed. - **Anything that must run inside your own cloud account** for compliance reasons. Runtime does not host your functions; it is where the model's code goes. Most teams end up with exactly that split: their front door and record-keeping on the function platform they already use, and every tool call in a sandbox. ## In short - A function runs your code with your credentials; a sandbox runs the model's code with none. Match the code to the box. - Keep the agent loop in a function only if a task fits the time limit; otherwise run it in the sandbox with its key as a secret it cannot read. - Call sandboxes from functions with `getOrCreate` and a stable name, so retries are safe and every turn finds its machine. - Most of a sandbox's cost is the time it stays awake; tune `idlePauseSeconds` to your turn length. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). A Runtime sandbox is ready for a request path: its first command runs 221 ms after the create request on Runtime's servers, and a paused one wakes for the next command in 153 ms. Read the [quick start](/docs/start) and call one from your existing function today.