Runtime

Best sandbox for LlamaIndex agents in 2026

LlamaIndex's code tool says it needs "heavy sandboxing or virtual machines"; the best sandbox is that VM, as a FunctionTool.

Runtime puts a LlamaIndex agent's code in a Firecracker microVM with its own kernel, wrapped as four ordinary FunctionTools, so a FunctionAgent or a workflow calls it like any other tool. The agent can read the documents your index retrieved, compute over them with pandas, and hand back a number it actually ran. Fifty thousand two-minute analysis calls a month, each busy for 20 CPU-seconds, cost $56.94 on Runtime against $276.00 on E2B and $176.89 on Vercel Sandbox, at rates checked 23 September to 2 October 2026.

Where does LlamaIndex run code today?

Tool Package Where the code runs
CodeInterpreterToolSpec llama-index-tools-code-interpreter 0.6.0 python -c in a subprocess on your own host
Azure code interpreter tool spec llama-index-tools-azure-code-interpreter Azure Container Apps dynamic sessions
Runtime sandbox_tools withruntime A Runtime microVM with a shell and files

The first tool's source, read on 1 October 2026, is plain about it: "Arbitrary code execution is possible on the machine running this tool", and it "is not recommended to be used in a production setting, and would require heavy sandboxing or virtual machines". It runs the model's code with subprocess.run beside your index, your keys and your data.

What should a sandbox give a retrieval agent?

  • A place to compute, apart from the index. Retrieval stays in your process; only the snippet the model wrote crosses into the sandbox, with the rows it needs written in as a file.
  • The data libraries already there. pandas, NumPy and matplotlib are in the default image, so a chart or a group-by needs no install step.
  • Files in, files out. sbx.files.write puts a CSV or a PDF in; sbx.files.read brings a chart back to show the user.
  • More than Python when it helps. The model can call jq, sqlite3 or gcc with the same runtime_exec tool.
  • A clean machine per conversation. One sandbox per chat keeps one user's uploads away from the next, and pauses when the chat goes quiet.

What does a month of LlamaIndex analysis cost?

Fifty thousand calls, each two minutes on a 2 vCPU, 4 GiB sandbox and busy for 20 CPU-seconds, at each provider's published rates:

Provider Isolation CPU billed on A month of calls
Runtime Firecracker microVM Measured use $56.94
Cloudflare Sandbox A container in its own VM Measured use $115.04
Vercel Sandbox Firecracker microVM Measured use $176.89
E2B Firecracker microVM Every vCPU held $276.00
Modal gVisor, a shared kernel The CPU requested $396.60

That month is 79% cheaper on Runtime than on E2B, because a call that spends most of its two minutes waiting for the model is billed for the CPU it uses, not the two vCPUs it holds.

A run on Runtime, recorded 1 October 2026

Terminalpip install withruntime llama-index-core llama-index-llms-anthropic "anthropic<1"
Pythonfrom llama_index.core.agent.workflow import FunctionAgentfrom llama_index.core.tools import FunctionToolfrom llama_index.llms.anthropic import Anthropicfrom withruntime import Sandboxfrom withruntime.tools import sandbox_toolsasync def ask(question: str) -> str:    with Sandbox.create() as sbx:        tools = [FunctionTool.from_defaults(fn=f) for f in sandbox_tools(sbx)]        agent = FunctionAgent(tools=tools, llm=Anthropic(model="claude-sonnet-4-6"))        return str(await agent.run(user_msg=question))

LlamaIndex has no scripted function-calling model to drive the agent without a paid call, so on 1 October 2026 we called the same FunctionTools directly, as the agent does, against a live sandbox with llama-index-core 0.14.25:

textruntime_write_file words.py   Wrote 139 bytes to /workspace/words.pyruntime_exec python3 words.py  exit_code 0  [('ubuntu', 9), ('url', 4), ('https', 4)]

The script counted words in the sandbox's own /etc/os-release, which is Ubuntu's, not your server's. Swapping the tool in an existing agent and mixing it with retrieval are in LlamaIndex in a sandbox.

When might another sandbox fit better?

  • Azure-only deployments. A team that must keep everything inside Azure can use the Azure dynamic sessions tool spec LlamaIndex ships.

Sources

Checked 1 October 2026.

Your first 100 hoursare on us.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Claim 100 hours free