Best sandbox for the OpenAI Agents SDK in 2026
For a SandboxAgent in production, pick a hosted client with its own kernel per sandbox, PTY sessions, and pause that keeps the workspace.
Runtime's RuntimeCloudSandboxClient drops into SandboxRunConfig beside
the seven hosted clients OpenAI documents, and runs the same agent for less
than any of them. Five thousand five-minute fix-the-test runs a month, each
busy for 45 CPU-seconds, cost $14.06 on Runtime,
$69.00 on E2B and $133.14 on
Runloop at the rates each published, checked 23 September to 2 October 2026. The agent
definition does not change; only the client in the run configuration does.
Which sandbox clients does the OpenAI Agents SDK support?
OpenAI's sandbox clients page,
read on 1 October 2026, lists two local clients and seven hosted ones, each an
extra of openai-agents:
| Client | Where commands run | Install |
|---|---|---|
UnixLocalSandboxClient |
Processes on your own host | Included |
DockerSandboxClient |
A container on your machine | openai-agents[docker] |
BlaxelSandboxClient |
Blaxel | openai-agents[blaxel] |
CloudflareSandboxClient |
Cloudflare | openai-agents[cloudflare] |
DaytonaSandboxClient |
Daytona | openai-agents[daytona] |
E2BSandboxClient |
E2B | openai-agents[e2b] |
ModalSandboxClient |
Modal | openai-agents[modal] |
RunloopSandboxClient |
Runloop | openai-agents[runloop] |
VercelSandboxClient |
Vercel | openai-agents[vercel] |
RuntimeCloudSandboxClient |
A Runtime Firecracker microVM | withruntime[openai-agents] |
OpenAI describes the local client as for trusted development: commands run as host processes. The hosted clients move "the workspace boundary to a provider-managed environment", which is where an agent that writes and runs its own code belongs.
What should a SandboxAgent's sandbox handle?
A SandboxAgent drives the sandbox through a small set of capabilities, and
each one is a place where a client can fall short:
- Commands that outlive a turn.
exec_commandreturns after its yield time; a build still running keeps a session thatwrite_stdinpolls. On Runtime these are real PTY sessions, so a REPL or an interactive installer takes typed input. - Edits by patch.
apply_patchchanges files of any size in place. - The manifest. Files, local directories, git repositories, environment variables, users and groups are staged before the first turn.
- Exposed ports. A port the agent's app listens on becomes a private HTTPS preview whose address carries its own token.
- Resuming.
workspace_persistence="snapshot"keeps the whole machine, memory and background processes included, andpause_on_exitwakes the very same sandbox on the next run.
How do the hosted clients compare on cost?
Each row prices the same month: 5,000 runs of a 2 vCPU, 4 GiB sandbox, five minutes long and busy for 45 CPU-seconds, at the provider's published rates.
| Provider | Isolation | CPU billed on | A month of runs | Runtime costs less by |
|---|---|---|---|---|
| Runtime | Firecracker microVM | Measured use | $14.06 | |
| Cloudflare Sandbox | A container in its own VM | Measured use | $28.26 | 50% |
| Vercel Sandbox | Firecracker microVM | Measured use | $43.33 | 68% |
| E2B | Firecracker microVM | Every vCPU held | $69.00 | 80% |
| Daytona | Containers by default | Every vCPU held | $69.00 | 80% |
| Blaxel | Lightweight VMs | Included in memory | $72.45 | 81% |
| Modal | gVisor, a shared kernel | The CPU requested | $99.15 | 86% |
| Runloop | A microVM, container inside | Every vCPU held | $133.14 | 89% |
A run like this spends most of its five minutes waiting for the model's next step. The providers that bill every vCPU held charge for that wait at full rate; Runtime bills it at its floor of a twentieth of a vCPU per vCPU.
A run on Runtime, recorded 1 October 2026
The code that runs the agent, with a real model:
Pythonfrom agents import Runnerfrom agents.run import RunConfigfrom agents.sandbox import SandboxAgent, SandboxRunConfigfrom withruntime.openai_agents import RuntimeCloudSandboxClient, RuntimeCloudSandboxClientOptionsagent = SandboxAgent(name="Coder", instructions="Fix the failing test, then run the suite.")config = RunConfig( sandbox=SandboxRunConfig( client=RuntimeCloudSandboxClient(), options=RuntimeCloudSandboxClientOptions(labels={"task": "fix-add"}), ))print(Runner.run_sync(agent, "The tests are in test_calc.py.", run_config=config).final_output)We ran the same agent with the SDK's TypeScript runner and its scripted test
model, ScriptedModel from @openai/agents/testing, so no model was paid for. The
manifest held a calc.py whose add subtracted, and a test for it. The
client created the sandbox, the runner made three tool calls, and the client
stopped the sandbox at the end:
textexec_command pip install --quiet pytest; python3 -m pytest -q FAILED test_calc.py::test_add - assert -1 == 5 1 failed in 0.02sapply_patch update_file calc.py: return a - b -> return a + b (completed)exec_command python3 -m pytest -q 1 passed in 0.00sfinal output Fixed add() and the suite passes.Every option of the client, from exposed_ports to snapshot_retention_days,
is in the OpenAI Agents SDK guide, and the
TypeScript version is in OpenAI Agents SDK in a sandbox.
Is the hosted Code Interpreter tool enough?
CodeInterpreterTool runs Python in a container OpenAI manages, and only for
OpenAIResponsesModel. It suits a model that needs a calculator. An agent that
clones a repository, installs packages, runs a test suite or serves an app
needs a machine your code holds, which is what a sandbox client gives it. The
two are compared in OpenAI Code Interpreter alternative.
When might another sandbox fit better?
- GPUs. Runtime sandboxes are CPU machines; a sandbox agent that needs a GPU fits Modal or Daytona.
Sources
Checked 1 October 2026.
- OpenAI Agents SDK: sandbox clients: the local and hosted clients, their extras, and the "provider-managed environment" description
- Each provider's published rates, checked 23 September to 2 October 2026, as the pricing guide lists them
- The recorded run: Runtime's TypeScript SDK with
@openai/agents0.18.0 and itsScriptedModel, against production on 1 October 2026