# Best sandbox for LlamaIndex agents in 2026 LlamaIndex's code tool says it needs "heavy sandboxing or virtual machines"; the best sandbox is that VM, as a `FunctionTool`. **Runtime puts a LlamaIndex agent's code in a Firecracker microVM with its own kernel, wrapped as four ordinary `FunctionTool`s, so a `FunctionAgent` or a workflow calls it like any other tool.** The agent can read the documents your index retrieved, compute over them with pandas, and hand back a number it actually ran. Fifty thousand two-minute analysis calls a month, each busy for 20 CPU-seconds, cost $56.94 on Runtime against $276.00 on E2B and $176.89 on Vercel Sandbox, at rates checked 23 September to 2 October 2026. ## Where does LlamaIndex run code today? | Tool | Package | Where the code runs | | -------------------------------- | ------------------------------------------ | -------------------------------------------- | | `CodeInterpreterToolSpec` | `llama-index-tools-code-interpreter` 0.6.0 | `python -c` in a subprocess on your own host | | Azure code interpreter tool spec | `llama-index-tools-azure-code-interpreter` | Azure Container Apps dynamic sessions | | Runtime `sandbox_tools` | `withruntime` | A Runtime microVM with a shell and files | The first tool's source, read on 1 October 2026, is plain about it: "Arbitrary code execution is possible on the machine running this tool", and it "is not recommended to be used in a production setting, and would require heavy sandboxing or virtual machines". It runs the model's code with `subprocess.run` beside your index, your keys and your data. ## What should a sandbox give a retrieval agent? - **A place to compute, apart from the index.** Retrieval stays in your process; only the snippet the model wrote crosses into the sandbox, with the rows it needs written in as a file. - **The data libraries already there.** pandas, NumPy and matplotlib are in the default image, so a chart or a group-by needs no install step. - **Files in, files out.** `sbx.files.write` puts a CSV or a PDF in; `sbx.files.read` brings a chart back to show the user. - **More than Python when it helps.** The model can call `jq`, `sqlite3` or `gcc` with the same `runtime_exec` tool. - **A clean machine per conversation.** One sandbox per chat keeps one user's uploads away from the next, and pauses when the chat goes quiet. ## What does a month of LlamaIndex analysis cost? Fifty thousand calls, each two minutes on a 2 vCPU, 4 GiB sandbox and busy for 20 CPU-seconds, at each provider's published rates: | Provider | Isolation | CPU billed on | A month of calls | | ------------------ | ------------------------- | ----------------- | ------------------------------------ | | Runtime | Firecracker microVM | Measured use | $56.94 | | Cloudflare Sandbox | A container in its own VM | Measured use | $115.04 | | Vercel Sandbox | Firecracker microVM | Measured use | $176.89 | | E2B | Firecracker microVM | Every vCPU held | $276.00 | | Modal | gVisor, a shared kernel | The CPU requested | $396.60 | That month is 79% cheaper on Runtime than on E2B, because a call that spends most of its two minutes waiting for the model is billed for the CPU it uses, not the two vCPUs it holds. ## A run on Runtime, recorded 1 October 2026 ```bash no-run pip install withruntime llama-index-core llama-index-llms-anthropic "anthropic<1" ``` ```python check from llama_index.core.agent.workflow import FunctionAgent from llama_index.core.tools import FunctionTool from llama_index.llms.anthropic import Anthropic from withruntime import Sandbox from withruntime.tools import sandbox_tools async def ask(question: str) -> str: with Sandbox.create() as sbx: tools = [FunctionTool.from_defaults(fn=f) for f in sandbox_tools(sbx)] agent = FunctionAgent(tools=tools, llm=Anthropic(model="claude-sonnet-4-6")) return str(await agent.run(user_msg=question)) ``` LlamaIndex has no scripted function-calling model to drive the agent without a paid call, so on 1 October 2026 we called the same `FunctionTool`s directly, as the agent does, against a live sandbox with llama-index-core 0.14.25: ```text runtime_write_file words.py Wrote 139 bytes to /workspace/words.py runtime_exec python3 words.py exit_code 0 [('ubuntu', 9), ('url', 4), ('https', 4)] ``` The script counted words in the sandbox's own `/etc/os-release`, which is Ubuntu's, not your server's. Swapping the tool in an existing agent and mixing it with retrieval are in [LlamaIndex in a sandbox](/integrations/llamaindex). ## When might another sandbox fit better? - **Azure-only deployments.** A team that must keep everything inside Azure can use the Azure dynamic sessions tool spec LlamaIndex ships. ## Sources Checked 1 October 2026. - [llama-index-tools-code-interpreter on PyPI](https://pypi.org/project/llama-index-tools-code-interpreter/): version 0.6.0, whose source holds the warnings quoted and the `subprocess.run` call - [llama-index-tools-azure-code-interpreter on PyPI](https://pypi.org/project/llama-index-tools-azure-code-interpreter/) - Each provider's published rates, checked 23 September to 2 October 2026, as the [pricing guide](/docs/pricing) lists them - The recorded run: Runtime's Python SDK with llama-index-core 0.14.25 on Python 3.12, against production on 1 October 2026