Best sandbox for LlamaIndex agents in 2026
LlamaIndex's code tool says it needs "heavy sandboxing or virtual machines"; the best sandbox is that VM, as a FunctionTool.
Runtime puts a LlamaIndex agent's code in a Firecracker microVM with its own
kernel, wrapped as four ordinary FunctionTools, so a FunctionAgent or a
workflow calls it like any other tool. The agent can read the documents your
index retrieved, compute over them with pandas, and hand back a number it
actually ran. Fifty thousand two-minute analysis calls a month, each busy for
20 CPU-seconds, cost $56.94 on Runtime against
$276.00 on E2B and $176.89 on
Vercel Sandbox, at rates checked 23 September to 2 October 2026.
Where does LlamaIndex run code today?
| Tool | Package | Where the code runs |
|---|---|---|
CodeInterpreterToolSpec |
llama-index-tools-code-interpreter 0.6.0 |
python -c in a subprocess on your own host |
| Azure code interpreter tool spec | llama-index-tools-azure-code-interpreter |
Azure Container Apps dynamic sessions |
Runtime sandbox_tools |
withruntime |
A Runtime microVM with a shell and files |
The first tool's source, read on 1 October 2026, is plain about it: "Arbitrary
code execution is possible on the machine running this tool", and it "is not
recommended to be used in a production setting, and would require heavy
sandboxing or virtual machines". It runs the model's code with subprocess.run
beside your index, your keys and your data.
What should a sandbox give a retrieval agent?
- A place to compute, apart from the index. Retrieval stays in your process; only the snippet the model wrote crosses into the sandbox, with the rows it needs written in as a file.
- The data libraries already there. pandas, NumPy and matplotlib are in the default image, so a chart or a group-by needs no install step.
- Files in, files out.
sbx.files.writeputs a CSV or a PDF in;sbx.files.readbrings a chart back to show the user. - More than Python when it helps. The model can call
jq,sqlite3orgccwith the sameruntime_exectool. - A clean machine per conversation. One sandbox per chat keeps one user's uploads away from the next, and pauses when the chat goes quiet.
What does a month of LlamaIndex analysis cost?
Fifty thousand calls, each two minutes on a 2 vCPU, 4 GiB sandbox and busy for 20 CPU-seconds, at each provider's published rates:
| Provider | Isolation | CPU billed on | A month of calls |
|---|---|---|---|
| Runtime | Firecracker microVM | Measured use | $56.94 |
| Cloudflare Sandbox | A container in its own VM | Measured use | $115.04 |
| Vercel Sandbox | Firecracker microVM | Measured use | $176.89 |
| E2B | Firecracker microVM | Every vCPU held | $276.00 |
| Modal | gVisor, a shared kernel | The CPU requested | $396.60 |
That month is 79% cheaper on Runtime than on E2B, because a call that spends most of its two minutes waiting for the model is billed for the CPU it uses, not the two vCPUs it holds.
A run on Runtime, recorded 1 October 2026
Terminalpip install withruntime llama-index-core llama-index-llms-anthropic "anthropic<1"Pythonfrom llama_index.core.agent.workflow import FunctionAgentfrom llama_index.core.tools import FunctionToolfrom llama_index.llms.anthropic import Anthropicfrom withruntime import Sandboxfrom withruntime.tools import sandbox_toolsasync def ask(question: str) -> str: with Sandbox.create() as sbx: tools = [FunctionTool.from_defaults(fn=f) for f in sandbox_tools(sbx)] agent = FunctionAgent(tools=tools, llm=Anthropic(model="claude-sonnet-4-6")) return str(await agent.run(user_msg=question))LlamaIndex has no scripted function-calling model to drive the agent without
a paid call, so on 1 October 2026 we called the same FunctionTools directly,
as the agent does, against a live sandbox with llama-index-core 0.14.25:
textruntime_write_file words.py Wrote 139 bytes to /workspace/words.pyruntime_exec python3 words.py exit_code 0 [('ubuntu', 9), ('url', 4), ('https', 4)]The script counted words in the sandbox's own /etc/os-release, which is
Ubuntu's, not your server's. Swapping the tool in an existing agent and
mixing it with retrieval are in LlamaIndex in a sandbox.
When might another sandbox fit better?
- Azure-only deployments. A team that must keep everything inside Azure can use the Azure dynamic sessions tool spec LlamaIndex ships.
Sources
Checked 1 October 2026.
- llama-index-tools-code-interpreter on PyPI:
version 0.6.0, whose source holds the warnings quoted and the
subprocess.runcall - llama-index-tools-azure-code-interpreter on PyPI
- Each provider's published rates, checked 23 September to 2 October 2026, as the pricing guide lists them
- The recorded run: Runtime's Python SDK with llama-index-core 0.14.25 on Python 3.12, against production on 1 October 2026