Best sandbox for the Claude Agent SDK in 2026
The best sandbox for the Claude Agent SDK starts in well under a second, keeps a session's memory while it waits, and bills only CPU used.
Runtime has an answer to each of the five questions Anthropic's hosting guide asks of a sandbox provider, and it costs the least of the sandboxes compared here. A thousand ten-minute agent sessions, each busy for one CPU-minute, cost $5.42 a month on Runtime against $27.60 on E2B and $39.66 on Modal, at rates checked 23 September to 2 October 2026. Each sandbox is a Firecracker microVM with its own Linux kernel, and a new one ran its first Python command 221 ms after the create request at the median on Runtime's servers (28 September 2026, speed).
What does the Claude Agent SDK need from a sandbox?
The SDK starts a claude subprocess that owns a shell, a working directory and
its session transcripts on local disk. Anthropic's
hosting guide, read on
1 October 2026, suggests 1 GiB of memory, 5 GiB of disk and one CPU per agent
as a starting point, and lists the questions to ask a provider. Here is how
Runtime answers each one:
| Anthropic's question | What it asks | Runtime |
|---|---|---|
| Who runs the sandbox | A service, or software you operate | A service: one API call, no cluster |
| Cold-start latency | Ephemeral sessions "need sub-second starts" | First command 221 ms after create, median |
| Persistent storage | Durable volumes, or only ephemeral disk | Volumes, snapshots, and pause that keeps memory for 1 to 365 days |
| Pricing model | "Per-second pricing suits bursty ephemeral" work | Per second, CPU billed on measured use |
| Networking | Egress rules, outbound proxies, private networks | Host-enforced allow-lists, secrets proxy, WireGuard tunnels |
Which sandbox fits each session pattern?
The guide names four session patterns. Runtime has a direct answer for each:
- Ephemeral, one sandbox per task.
await using sbx = await Sandbox.create()starts it and stops it when the task returns, even after an error. - Long-running, an agent that serves traffic for days. A paid sandbox runs for as long as its lease is extended, with no session cap; a preview gives its port an HTTPS address.
- Hybrid, a session that sits idle between visits. A Runtime sandbox
pauses itself after 60 seconds with nothing happening and wakes on the
next call with its memory intact. When the SDK runs inside the sandbox, its
claudeprocess is still running on wake, so the conversation resumes without first reloading a transcript store. - Multi-agent, several SDK processes in one machine. Size the sandbox up to 16 vCPUs and 64 GiB; idle vCPUs cost nothing, because CPU is billed on use.
How do the sandboxes compare for Agent SDK sessions?
The workload is a thousand sessions of 2 vCPU and 4 GiB, ten minutes each and busy for 60 CPU-seconds, at each provider's published rates.
| Provider | Isolation | CPU billed on | A month of sessions |
|---|---|---|---|
| Runtime | Firecracker microVM | Measured use | $5.42 |
| E2B | Firecracker microVM | Every vCPU held | $27.60 |
| Daytona | Containers by default | Every vCPU held | $27.60 |
| Modal | gVisor, a shared kernel | The CPU requested | $39.66 |
| Vercel Sandbox | Firecracker microVM | Measured use | $16.27 |
| Cloudflare Sandbox | A container in its own VM | Measured use | $10.70 |
Runtime costs 80% less than E2B here and 49% less than Cloudflare, the cheapest of the others. The gap comes from what the session does: most of its ten minutes are spent waiting on Claude, and Runtime bills that wait at its CPU floor of a twentieth of a vCPU rather than for every vCPU held.
A run on Runtime, recorded 1 October 2026
runtimeMcpServer(sbx) from withruntime/claude-agent-sdk is the in-process
MCP server that query() is handed. This is the code an application runs:
TypeScriptimport { query } from "@anthropic-ai/claude-agent-sdk";import { Sandbox } from "withruntime";import { RUNTIME_TOOL_NAMES, runtimeMcpServer } from "withruntime/claude-agent-sdk";await using sbx = await Sandbox.create();for await (const message of query({ prompt: "Write primes.py that lists the primes below 60, run it, and say which kernel ran it.", options: { mcpServers: { runtime: runtimeMcpServer(sbx) }, allowedTools: RUNTIME_TOOL_NAMES, disallowedTools: ["Bash", "Read", "Write", "Edit"], },})) { if (message.type === "result" && message.subtype === "success") console.log(message.result);}We called the same server's tools over MCP against a live sandbox, as the SDK's subprocess would, without a paid model call. What came back:
texttools: runtime_exec, runtime_read_file, runtime_write_file, runtime_list_filesruntime_write_file primes.py -> Wrote 68 bytes to /workspace/primes.pyruntime_exec "python3 primes.py && uname -sr" exitCode 0 [2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37, 41, 43, 47, 53, 59] Linux 6.1.186The kernel line is the sandbox's own: a microVM boots its own Linux, not the host's. The full integration, with the Python SDK and spending caps, is Claude Agent SDK in a sandbox.
Should the whole SDK run inside the sandbox instead?
Both shapes work. Keeping query() in your service and handing it the MCP
server means the Anthropic key never enters the sandbox, and the sandbox only
sees commands. Running the SDK inside the sandbox, as Anthropic's container
pattern does, suits a per-user agent that should own its whole machine: store
the key as a Runtime secret, and the sandbox holds only a placeholder that the
host swaps for the real key on requests to api.anthropic.com
(secrets). Anthropic's own hosted
option, Claude Managed Agents, can also be given
Runtime sandboxes through the remote MCP server.
When might another sandbox fit better?
- GPUs. Runtime sandboxes are CPU machines. An agent that trains or serves a model on a GPU fits Modal or Daytona.
Sources
Checked 1 October 2026.
- Hosting the Agent SDK: the subprocess model, the four session patterns, the provider questions and the 1 GiB, 5 GiB, one CPU starting point
- Securely deploying AI agents: the isolation options and the credential proxy pattern
- Each provider's published rates, checked 23 September to 2 October 2026, as the pricing guide lists them
- The recorded run: Runtime's TypeScript SDK with
@anthropic-ai/claude-agent-sdk0.3.277, against production on 1 October 2026