Build a code interpreter tool for any LLM in under 100 lines
A code interpreter tool is one function the model can call: it runs Python in a sandbox and returns text and charts the model can read.
Runtime (withruntime.com) runs each cell in a Firecracker microVM where variables persist between calls, charts come back as PNG images, and a command in a running sandbox takes 53 ms on Runtime's servers. This post builds the tool from scratch: first the obvious version and the way it fails, then the version that holds up, with a complete agent loop in under 100 lines and a table of what each model API expects.
What is a code interpreter tool, exactly?
It is a tool definition the model sees, plus a function you run when the model calls it. The model writes Python, your code runs it somewhere safe, and the output goes back into the conversation as the tool's result.
Three parts decide whether it works well:
- The description. It tells the model what is installed, whether state
persists, and how to get output back. A model that does not know
pandasis there will try to install it. - The runner. Where the code runs and what it can reach. Model-written code is untrusted code, even when your user means no harm.
- The result format. What the model reads back. This is the part most first versions get wrong, and the part that decides how often the model recovers from its own mistakes.
The code interpreter glossary entry covers the idea in general. The rest of this post is the engineering.
Why does the obvious version fail?
The obvious version writes the code to a file and runs it with python3. It
is ten lines, and it breaks on the second call:
TypeScriptimport { Sandbox } from "withruntime";// Version one: every call is a fresh Python process.async function runPython(sbx: Sandbox, code: string): Promise<string> { await sbx.files.write("/workspace/cell.py", code); const run = await sbx.exec("python3 cell.py", { cwd: "/workspace", timeoutMs: 30_000 }); if (run.timedOut) return "[timeout] Stopped after 30 s."; return `${run.stdout}${run.stderr}`.trim() || "[no output]";}await using sbx = await Sandbox.create({ network: { internet: false } });console.log( await runPython(sbx, "import pandas as pd\ndf = pd.DataFrame({'x': [1, 2, 3]})\nprint(len(df))"),);console.log(await runPython(sbx, "print(df['x'].sum())"));Pythonfrom withruntime import Sandbox# Version one: every call is a fresh Python process.def run_python(sbx: Sandbox, code: str) -> str: sbx.files.write("/workspace/cell.py", code) run = sbx.exec("python3 cell.py", cwd="/workspace", timeout_ms=30_000) if run.timed_out: return "[timeout] Stopped after 30 s." return (run.stdout + run.stderr).strip() or "[no output]"with Sandbox.create(network={"internet": False}) as sbx: print(run_python(sbx, "import pandas as pd\ndf = pd.DataFrame({'x': [1, 2, 3]})\nprint(len(df))")) print(run_python(sbx, "print(df['x'].sum())"))The first call prints 3. The second prints
NameError: name 'df' is not defined, because df died with the first
process.
This is not a contrived test. Models write analysis code the way people write
notebooks: load the data in one call, look at it in the next, chart it in a
third. Tell a model that nothing persists and it will reload the CSV in every
call, which costs time and tokens. Forget to tell it, and it spends turns
chasing NameError. Either way you pay for the gap between how the model
thinks and how the tool behaves.
The fix is a stateful session: one long-lived Python process per conversation
that runs each call as a cell. On Runtime that is sbx.interpreter.run(code).
Variables, imports and loaded files stay in memory between calls, and figures
still open when a cell ends come back as PNG images.
What should the tool return when the code goes wrong?
It should return the failure as an ordinary result, written so the model can act on it. A tool that throws ends the turn; a tool that returns a clear error lets the model fix its own code, which it does well.
| What the code does | What the cell returns | What the model should read |
|---|---|---|
| Raises an exception | status: "error", with name and traceback |
The error and the last lines of the traceback |
| Loops forever | status: "timeout" after timeoutMs |
That it stopped, and that earlier variables survive |
| Prints a whole DataFrame | Megabytes of stdout |
The head and the tail, with a count of what was cut |
| Draws a matplotlib chart | results[].data["image/png"], as base64 |
The image itself, if the model can read images |
| Ends with an expression | results[].data["text/plain"] |
The value, as a notebook would show it |
| Kills its own process | status: "lost", error ContextDied |
That the session restarted and its variables are gone |
| Reaches for the internet | A connection error from inside the sandbox | Nothing special: the error already says it |
| Prints nothing | Empty stdout |
[no output], so silence is not mistaken for a hang |
Four details in that table matter more than they look:
- Cut from the middle. Long output starts with the useful header and ends with the error or the total. Keep both ends and say how much you cut.
- Strip terminal colors. Python tracebacks from an interactive session can carry color codes. They waste tokens and confuse some models.
- Keep the last lines of a traceback. The bottom frames name the line that failed; the top frames are the interpreter's own machinery.
- A timeout keeps the session. On Runtime a timed-out cell is a result
with
status: "timeout", not an exception, and the context keeps its variables. Say so, or the model will reload everything.
What does the whole program look like?
Here is the complete loop: the tool definition, the stateful runner with every row of that table, and the model calls, under 100 lines in each language. It calls Anthropic's Messages API over plain HTTP so you can see every field; the next section maps the same fields to other APIs.
TypeScriptimport { Sandbox } from "withruntime";type Block = Record<string, unknown>;const MODEL = "claude-opus-5-5";const LIMIT = 4_000; // characters of output the model reads per callconst COLORS = /\x1b\[[0-9;]*m/g; // terminal color codes in tracebacksconst TOOL = { name: "run_python", description: "Run Python in a persistent session, like a notebook cell. Variables, imports and files " + "persist between calls. pandas, NumPy and matplotlib are installed; there is no internet. " + "Print what you need to see. Open matplotlib figures come back as images.", input_schema: { type: "object", properties: { code: { type: "string", description: "The Python code to run" } }, required: ["code"], },};function clip(text: string): string { if (text.length <= LIMIT) return text; const cut = text.length - LIMIT; return `${text.slice(0, LIMIT / 2)}\n[... ${cut} characters cut ...]\n${text.slice(-LIMIT / 2)}`;}async function runPython(sbx: Sandbox, code: string): Promise<Block[]> { const cell = await sbx.interpreter.run(code, { timeoutMs: 30_000 }); const text: string[] = []; const images: Block[] = []; if (cell.stdout) text.push(cell.stdout); if (cell.stderr) text.push(`[stderr]\n${cell.stderr}`); for (const result of cell.results) { const ref = result.refs["image/png"]; const inline = result.data["image/png"]; const value = result.data["text/plain"]; const data = ref ? Buffer.from(await sbx.interpreter.result(ref)).toString("base64") : inline; if (typeof data === "string") images.push({ type: "image", source: { type: "base64", media_type: "image/png", data } }); else if (typeof value === "string") text.push(value); } if (cell.error) { const trace = cell.error.traceback.replace(COLORS, "").split("\n"); text.push(`[${cell.error.name}] ${cell.error.value}\n${trace.slice(-12).join("\n")}`); } if (cell.status === "timeout") text.push("[timeout] Stopped after 30 s; variables are kept."); if (cell.status === "lost") text.push("[restarted] The session stopped; variables are gone."); return [{ type: "text", text: clip(text.join("\n").trim() || "[no output]") }, ...images];}async function ask(messages: Block[]): Promise<Block> { const response = await fetch("https://api.anthropic.com/v1/messages", { method: "POST", headers: { "content-type": "application/json", "x-api-key": process.env.ANTHROPIC_API_KEY ?? "", "anthropic-version": "2023-06-01", }, body: JSON.stringify({ model: MODEL, max_tokens: 16_000, tools: [TOOL], messages }), }); if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`); return (await response.json()) as Block;}async function chat(question: string): Promise<string> { await using sbx = await Sandbox.create({ memoryMiB: 2048, network: { internet: false } }); const messages: Block[] = [{ role: "user", content: question }]; for (let turn = 0; turn < 20; turn++) { const reply = await ask(messages); const content = reply.content as Block[]; messages.push({ role: "assistant", content }); if (reply.stop_reason !== "tool_use") { const texts = content.filter((block) => block.type === "text"); return texts.map((block) => block.text).join("\n"); } const results: Block[] = []; for (const call of content.filter((block) => block.type === "tool_use")) { const output = await runPython(sbx, (call.input as { code: string }).code); results.push({ type: "tool_result", tool_use_id: call.id, content: output }); } messages.push({ role: "user", content: results }); } return "Stopped after 20 model turns.";}console.log(await chat("Roll two dice 10,000 times and chart how often each sum comes up."));Pythonimport base64import jsonimport osimport reimport urllib.requestfrom withruntime import SandboxMODEL = "claude-opus-5-5"LIMIT = 4_000 # characters of output the model reads per callTOOL = { "name": "run_python", "description": ( "Run Python in a persistent session, like a notebook cell. Variables, imports and files " "persist between calls. pandas, NumPy and matplotlib are installed; there is no internet. " "Print what you need to see. Open matplotlib figures come back as images." ), "input_schema": { "type": "object", "properties": {"code": {"type": "string", "description": "The Python code to run"}}, "required": ["code"], },}def clip(text: str) -> str: if len(text) <= LIMIT: return text return f"{text[:LIMIT // 2]}\n[... {len(text) - LIMIT} characters cut ...]\n{text[-LIMIT // 2:]}"def run_python(sbx: Sandbox, code: str) -> list[dict]: cell = sbx.interpreter.run(code, timeout_ms=30_000) text, images = [], [] if cell["stdout"]: text.append(cell["stdout"]) if cell["stderr"]: text.append(f"[stderr]\n{cell['stderr']}") for result in cell["results"]: ref = result["refs"].get("image/png") png = base64.b64encode(sbx.interpreter.result(ref)).decode() if ref else result["data"].get("image/png") if png: images.append({"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": png}}) elif "text/plain" in result["data"]: text.append(str(result["data"]["text/plain"])) if cell["error"]: error = cell["error"] trace = re.sub(r"\x1b\[[0-9;]*m", "", error["traceback"]).splitlines()[-12:] text.append(f"[{error['name']}] {error['value']}\n" + "\n".join(trace)) if cell["status"] == "timeout": text.append("[timeout] Stopped after 30 s; variables are kept.") if cell["status"] == "lost": text.append("[restarted] The session stopped; variables are gone.") return [{"type": "text", "text": clip("\n".join(text).strip() or "[no output]")}, *images]def ask(messages: list[dict]) -> dict: request = urllib.request.Request( "https://api.anthropic.com/v1/messages", data=json.dumps({"model": MODEL, "max_tokens": 16_000, "tools": [TOOL], "messages": messages}).encode(), headers={ "content-type": "application/json", "x-api-key": os.environ["ANTHROPIC_API_KEY"], "anthropic-version": "2023-06-01", }, ) with urllib.request.urlopen(request, timeout=600) as response: return json.load(response)def chat(question: str) -> str: with Sandbox.create(memory_mib=2048, network={"internet": False}) as sbx: messages: list[dict] = [{"role": "user", "content": question}] for _ in range(20): reply = ask(messages) messages.append({"role": "assistant", "content": reply["content"]}) if reply["stop_reason"] != "tool_use": return "\n".join(block["text"] for block in reply["content"] if block["type"] == "text") results = [ {"type": "tool_result", "tool_use_id": call["id"], "content": run_python(sbx, call["input"]["code"])} for call in reply["content"] if call["type"] == "tool_use" ] messages.append({"role": "user", "content": results}) return "Stopped after 20 model turns."print(chat("Roll two dice 10,000 times and chart how often each sum comes up."))A few choices in there are deliberate:
- The assistant's content goes back whole. The loop appends the model's entire reply, not just its text, so tool calls and any reasoning blocks stay in the history the API expects.
- Every tool call gets a result. When the model asks for two cells at once, both results go back in one user message, in any order, matched by id.
- The loop has a ceiling. Twenty model turns is generous for one question. Without a ceiling, a model stuck on an error can loop until your budget notices.
- The chart goes to the model. Returning the PNG lets the model look at its own chart and fix an unreadable axis before your user sees it.
How do you plug it into other model APIs?
You change the tool's wrapper and the names of three fields; the runner stays
the same. Every major API uses JSON Schema for the arguments, so the code
property carries over unchanged.
| API | Tool definition | The model asks to run code | You send the result back |
|---|---|---|---|
| Anthropic Messages | { name, description, input_schema } |
A tool_use block with id and input |
A tool_result block with tool_use_id, in a user message |
| OpenAI Chat Completions | { type: "function", function: { name, description, parameters } } |
tool_calls[], arguments as a JSON string |
A message with role: "tool" and tool_call_id |
| OpenAI Responses | { type: "function", name, description, parameters } |
A function_call item with call_id and arguments |
A function_call_output item with the same call_id |
| Gemini | functionDeclarations: [{ name, description, parameters }] |
A functionCall part with name and args |
A functionResponse part with name and response |
Two things trip people up when they port the loop. OpenAI's arguments arrive
as a JSON string, so parse them before reading code, and expect the parse to
fail now and then. And not every API accepts an image inside a tool result.
Where it does not, return the text and send the chart in a user message right
after it.
If you would rather not write the loop at all, Runtime's
withruntime/tools package ships ready-made tools for commands and files,
each a name, a description and a JSON Schema you wrap the same way (an LLM with a code execution tool),
and an agent connected over MCP gets the interpreter as
runtime_sandbox_interpreter_run.
What should the sandbox be allowed to reach?
Give the interpreter the data it needs and nothing else. The code comes from a model, and the model reads whatever your user uploads, so treat each cell as code from a stranger.
- No internet by default. The program above creates the sandbox with
network: { internet: false }. The default image already has NumPy, pandas and matplotlib, so most analysis needs nothing from outside (turn off a sandbox's internet). - Allow a package index if you must. If the model needs a library the image
lacks, allow
pypi.organdfiles.pythonhosted.organd nothing more. Apip installinside a cell is importable in that same session. - No keys in the sandbox. The interpreter needs none. Your model key stays in your server process, which calls the model; the sandbox only runs cells.
- Files go in, not credentials. Put the user's CSV in with
sbx.files.write("/workspace/data.csv", bytes)and tell the model the path.
Why each of these matters, step by step, is in how prompt injection becomes code execution.
How long should one sandbox live?
One sandbox per conversation, paused between messages. The program above deletes its sandbox when the question is answered, which is right for one-shot questions. For a chat, look the sandbox up by a name built from the conversation id, so the user's DataFrame is still in memory when the next message arrives. A paused sandbox keeps its interpreter's variables and wakes 76 ms after the wake request, and any interpreter call wakes it by itself (pause, don't delete).
The machine is the cheap part. Take an analysis conversation with six cells of 3 CPU-seconds each, four minutes from first to last message, then 60 seconds of quiet before it pauses itself, with 2 GiB of memory:
TextMemory: (240 + 60) s / 3,600 × 2 GiB × $0.0075 = $0.001250CPU: 6 × 3 / 3,600 × $0.025 = $0.000125Total: $0.0013751,000 conversations: $1.3750The model tokens for those six cells will cost more than the sandbox that ran them. Size the memory to the data your users upload, since memory is the larger line (size CPU and memory).
In short
- A code interpreter tool is a description, a runner and a result format; the result format decides how often the model recovers.
- Run each conversation's code in one stateful session, because models write code like notebooks.
- Return errors, timeouts and lost sessions as results the model can act on, and cut long output from the middle.
- The same runner fits every major model API; only the tool wrapper and three field names change.
- Keep the sandbox offline, keyless and paused between messages.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Every Runtime sandbox has a stateful interpreter for Python and six other languages, runs a command 53 ms after the request on Runtime's servers, and bills $0.025 per vCPU-hour of CPU actually used. Start with 100 free hours, no card: sign in, or read get started and the code interpreter for chatbots guide.