Claude tool use: how to run the code Claude writes in a sandbox
Give Claude a run_command tool, run each tool_use block in a microVM, and reply with a tool_result holding the output and is_error.
A Runtime sandbox gives Claude a whole Linux machine for the conversation,
with the internet off and billed on the CPU it uses. Every sandbox is a
Firecracker microVM with its own kernel, running Ubuntu 24.04 with Python 3.12,
Node.js 24, git and gcc. Waiting between Claude's turns, a 2 vCPU, 4 GiB
sandbox costs $0.03125 an hour (pricing). The code uses
claude-opus-5-5, the model Anthropic's overview told developers to start
with on 25 September 2026.
One sandbox for the whole conversation
Claude often works in steps: write a file, run it, read the error, fix it. So this loop creates one sandbox before the first request and keeps it until Claude stops asking for the tool. Files from one call are there for the next.
TypeScriptimport Anthropic from "@anthropic-ai/sdk";import { Sandbox } from "withruntime";const anthropic = new Anthropic(); // ANTHROPIC_API_KEY stays in this processconst tools: Anthropic.Tool[] = [ { name: "run_command", description: "Run one bash command in a private Ubuntu 24.04 sandbox that has Python 3.12, " + "Node.js 24 and gcc but no internet. The working directory is /workspace and " + "files persist between calls in this conversation. Returns the exit code, " + "stdout and stderr; each command is stopped after 60 seconds.", input_schema: { type: "object", properties: { command: { type: "string", description: "A bash command line" } }, required: ["command"], }, },];await using sbx = await Sandbox.create({ network: { internet: false }, timeoutSeconds: 1800, onLeaseEnd: "stop",});const messages: Anthropic.MessageParam[] = [ { role: "user", content: "Write a prime sieve in C, compile it and print the primes below 100." },];for (let turn = 0; turn < 10; turn++) { const reply = await anthropic.messages.create({ model: "claude-opus-5-5", max_tokens: 4096, tools, messages, }); messages.push({ role: "assistant", content: reply.content }); if (reply.stop_reason !== "tool_use") { for (const block of reply.content) if (block.type === "text") console.log(block.text); break; } const results: Anthropic.ToolResultBlockParam[] = []; for (const block of reply.content) { if (block.type !== "tool_use") continue; const { command } = block.input as { command: string }; const run = await sbx.exec(["bash", "-c", command], { timeoutMs: 60_000 }); results.push({ type: "tool_result", tool_use_id: block.id, content: JSON.stringify({ exit_code: run.exitCode, timed_out: run.timedOut, stdout: run.stdout, stderr: run.stderr, }), is_error: run.exitCode !== 0, }); } messages.push({ role: "user", content: results });}Pythonimport jsonimport anthropicfrom withruntime import Sandboxclient = anthropic.Anthropic() # ANTHROPIC_API_KEY stays in this processtools = [ { "name": "run_command", "description": "Run one bash command in a private Ubuntu 24.04 sandbox that has Python 3.12, " "Node.js 24 and gcc but no internet. The working directory is /workspace and " "files persist between calls in this conversation. Returns the exit code, " "stdout and stderr; each command is stopped after 60 seconds.", "input_schema": { "type": "object", "properties": {"command": {"type": "string", "description": "A bash command line"}}, "required": ["command"], }, }]with Sandbox.create(network={"internet": False}, timeout_seconds=1800, on_lease_end="stop") as sbx: messages = [{"role": "user", "content": "Write a prime sieve in C, compile it and print the primes below 100."}] for _ in range(10): reply = client.messages.create(model="claude-opus-5-5", max_tokens=4096, tools=tools, messages=messages) messages.append({"role": "assistant", "content": reply.content}) if reply.stop_reason != "tool_use": print("".join(block.text for block in reply.content if block.type == "text")) break results = [] for block in reply.content: if block.type != "tool_use": continue run = sbx.exec(["bash", "-c", block.input["command"]], timeout_ms=60_000) output = {"exit_code": run.exit_code, "timed_out": run.timed_out, "stdout": run.stdout, "stderr": run.stderr} results.append( {"type": "tool_result", "tool_use_id": block.id, "content": json.dumps(output), "is_error": run.exit_code != 0} ) messages.append({"role": "user", "content": results})The command is Claude's to write, so it runs under bash -c inside the
sandbox, where a shell is exactly what Claude asked for. Your own process never
hands the text to a shell. The ten-turn cap bounds the loop, and leaving the
block stops the sandbox.
The rules Anthropic sets for tool results
Anthropic's guide on handling tool calls, read 25 September 2026, is strict about shape. A malformed reply is a 400 error, not a wrong answer.
| Rule | How the loop meets it |
|---|---|
stop_reason is tool_use when Claude wants a client tool |
The loop runs tools only then |
Each result names the id of its tool_use block |
tool_use_id: block.id |
Results come in the very next user message, first in content |
Every result goes in one user message, nothing else |
is_error: true marks a failed execution |
Set from a non-zero exit code |
| Results can carry text the model should not obey | Output stays inside tool_result, never in the prompt |
Claude may call the tool more than once in a turn; every tool_use block gets
its own result. With is_error set, Claude reads the traceback and tries
again, which is the point of returning stderr whole.
Anthropic's code execution tool, or this?
Anthropic runs its own server tool, which "allows Claude to run Bash commands and manipulate files" in a container Anthropic hosts. Its page, checked 25 September 2026, gives the container 1 CPU, 5 GiB of memory, and internet "completely disabled for security", so Claude "can't download or install additional packages at runtime". Past 1,550 free hours a month per organization it costs $0.05 an hour per container, with a 5-minute minimum.
A Runtime sandbox differs where a build or an analysis gets stuck:
- Packages on your terms. Allow PyPI or npm for an install, then turn the internet off before Claude's code runs; rules apply at once and bind root (the network).
- Size you choose. 2 vCPUs and 4 GiB by default, and
vcpu, memory and disk are set per sandbox. - State that outlives a request. A paused sandbox keeps its files, memory and processes for 1 to 365 days.
- Any model. The same tool serves Claude and every other model you call.
Costs for short and long runs side by side are in Claude code execution tool alternative.
Where the Anthropic key goes
Nowhere near the sandbox. The loop calls the Messages API from your server, and
only Claude's commands run inside, so ANTHROPIC_API_KEY stays in your
process. When Claude Code or another agent runs inside the sandbox and calls
Anthropic itself, store the key as a Runtime secret bound to
api.anthropic.com; the sandbox then holds only a placeholder that is
worthless anywhere else (secrets).
Claude Code in a sandbox walks through that case.
Using the Claude Agent SDK
The Agent SDK runs its own loop. runtimeMcpServer(sbx) from
withruntime/claude-agent-sdk hands it Runtime's command and file tools bound
to one sandbox (Claude Agent SDK).
Start on the free trial
New accounts get 50 free sandbox hours, no card. The first command prints a link to approve in your browser:
Terminalnpx withruntime sandbox run --trial -- gcc --versionMore: run untrusted LLM code, coding agent sandbox, egress control.
Sources
Checked 25 September 2026.
- Anthropic: define tools
- Anthropic: handle tool calls
- Anthropic: models overview
- Anthropic: code execution tool
- @anthropic-ai/sdk on npm (0.128.0) and anthropic on PyPI (1.8.0)
Facts on this page were checked on 25 September 2026.