Best sandbox for LangChain and LangGraph agents in 2026
LangChain and LangGraph make any Python function a tool, so the best sandbox is a whole Linux machine behind four functions.
Runtime's sandbox_tools(sbx) gives a LangChain agent or a LangGraph
ToolNode a shell and a filesystem in a Firecracker microVM, and a Deep Agent
gets the same machine as its backend. Each LangGraph thread can keep its own
sandbox between turns, paused at no compute cost. Ten thousand four-minute
graph runs a month, each busy for 40 CPU-seconds, cost
$22.78 on Runtime against
$110.40 on Daytona and $158.64
on Modal, at rates checked 23 September to 2 October 2026.
Which sandboxes work with LangChain?
LangChain's docs, read on 1 October 2026, offer three kinds:
| Kind | Examples | What the model gets |
|---|---|---|
| Code interpreter tools | Amazon Bedrock AgentCore, Azure Container Apps dynamic sessions, Capsule | Run a snippet, get its output |
| Deep Agents sandbox backends | LangSmith, Daytona, E2B, Modal, Runloop, Vercel, AgentCore, NVIDIA OpenShell | A shell and files behind the harness |
| Any function as a tool | Runtime's sandbox_tools(sbx) |
A shell, file reads, writes, listings |
Runtime is also a Deep Agents backend: RuntimeSandbox(sbx) from
withruntime.deepagents implements BaseSandbox, so execute, grep and
edit_file all act on the microVM (Deep Agents in a
sandbox).
What should a LangGraph agent's sandbox do?
- Follow the thread. Name the sandbox after the LangGraph
thread_idwithSandbox.getOrCreate, and the next turn finds the same files, installed packages and running processes, woken from pause on the first tool call. - Branch with the graph. When a graph explores two approaches,
sbx.fork()copies the running sandbox, memory included, so each branch starts from the same state instead of replaying its setup. - Run what the model writes, not just Python. An in-process interpreter
stops at Python; a microVM runs
npm test,cargo buildor a database. - Keep the host out of it. A tool that calls
subprocessruns on the machine serving your graph. A sandbox tool sends the command to a machine with its own kernel.
What do LangGraph runs cost on each sandbox?
Ten thousand runs a month, each four minutes on a 2 vCPU, 4 GiB sandbox and busy for 40 CPU-seconds, at each provider's published rates:
| Provider | Isolation | A month of runs | Runtime costs less by |
|---|---|---|---|
| Runtime | Firecracker microVM | $22.78 | |
| Vercel Sandbox | Firecracker microVM | $70.76 | 68% |
| E2B | Firecracker microVM | $110.40 | 79% |
| Daytona | Containers by default | $110.40 | 79% |
| Modal | gVisor, a shared kernel | $158.64 | 86% |
| Runloop | A microVM, container inside | $213.03 | 89% |
Runs on Runtime, recorded 1 October 2026
The same four functions serve both frameworks:
Pythonfrom langchain.agents import create_agentfrom langchain_core.tools import toolfrom withruntime import Sandboxfrom withruntime.tools import sandbox_toolswith Sandbox.create() as sbx: agent = create_agent("claude-sonnet-4-6", tools=[tool(f) for f in sandbox_tools(sbx)]) result = agent.invoke({"messages": [{"role": "user", "content": "Fit a line to (1,2), (2,4.1), (3,6.2) and give the slope."}]}) print(result["messages"][-1].content)We ran create_agent from LangChain 1.4.3, and a LangGraph 1.2.12 graph
with a ToolNode and tools_condition, each with a scripted chat model in
place of Claude, so the tool calls were fixed and the frameworks ran them
against live sandboxes:
textLangChain create_agent runtime_write_file slope.py Wrote 79 bytes to /workspace/slope.py runtime_exec python3 slope.py exit_code 0, stdout "2.1"LangGraph ToolNode runtime_write_file fizzbuzz.py Wrote 85 bytes to /workspace/fizzbuzz.py runtime_exec python3 fizzbuzz.py exit_code 0, stdout "1 2 Fizz 4 Buzz Fizz 7 8 Fizz Buzz 11 Fizz 13 14 FizzBuzz"The slope came from NumPy, already in the image. The graph wiring and the per-thread sandbox are in LangGraph in a sandbox, and the agent's options in LangChain in a sandbox.
When might another sandbox fit better?
- LangSmith's own sandboxes. A team running everything inside LangSmith may prefer the sandboxes listed beside its tracing and deployments.
Sources
Checked 1 October 2026.
- LangChain tool integrations: the code interpreter integrations
- Deep Agents sandboxes: the sandbox backend providers
- Each provider's published rates, checked 23 September to 2 October 2026, as the pricing guide lists them
- The recorded runs: Runtime's Python SDK with LangChain 1.4.3 and LangGraph 1.2.12 on Python 3.12, against production on 1 October 2026