# Build a coding agent on an open model such as DeepSeek or Qwen A coding agent is a loop: send the task and four tools to any OpenAI-compatible model, run each tool call in a sandbox, stop on finish. **Runtime (withruntime.com) is where the tool calls run: each task gets a Firecracker microVM that runs its first command 221 ms after the create request on Runtime's servers, and bills CPU at $0.025 per vCPU-hour of actual use while the model thinks.** Open-weight models such as DeepSeek and Qwen now write code well enough to fix real bugs, at a fraction of the price per token of the largest closed models. What they need is the same thing every coding agent needs: a machine where a wrong command costs nothing. This post builds the whole agent, explains the lines that make open models reliable, and grades the result without trusting the model's word. ## Which models and endpoints does this work with? Any endpoint that speaks the OpenAI chat completions API with `tools`. The loop below reads three environment variables, so switching models is a configuration change: | Where the model runs | `LLM_BASE_URL` | `LLM_MODEL`, for example | | ---------------------- | -------------------------------------------------------- | ------------------------- | | DeepSeek's API | `https://api.deepseek.com` | `deepseek-flash` | | Qwen on Alibaba Cloud | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` | `qwen3-coder-plus` | | OpenRouter | `https://openrouter.ai/api/v1` | any model listing `tools` | | Your own GPU with vLLM | `http://your-gpu-host:8000/v1` | the name you served | Model names change often; check the provider's model list before you copy one. Provider details that break loops, such as DeepSeek requiring its reasoning text sent back, are on the [DeepSeek](/integrations/deepseek-api) and [OpenRouter](/integrations/openrouter) pages. The loop below handles that case by appending the whole assistant message. ## What tools does a coding agent need? Four. Every extra tool is one more way for a smaller model to pick the wrong one, so keep the set small and make each tool forgiving: | Tool | What it does | Why this shape | | ----------- | ----------------------------------------------------- | ---------------------------------------------------------------- | | `run` | A bash command in the repo, with exit code and output | One tool covers tests, `grep`, `ls`, installs and `git` | | `read_file` | A file's text | Saves the model from quoting `cat` and losing a file to escaping | | `edit_file` | Replace one exact snippet, or write a new file | Small edits fail loudly instead of rewriting a file wrongly | | `finish` | Ends the task with a summary | A stop signal you can see, instead of guessing from silence | `edit_file` deserves the most care. Asking a model to rewrite a whole file invites it to drop a function it did not mean to touch. A replace of one exact snippet that refuses when the snippet appears zero or several times gives the model a precise error it can act on. ## What does the whole agent look like? Under a hundred lines in each language. The sandbox is created once per task and kept for every turn, so installed packages, built files and the repo's state carry from one tool call to the next: ```ts check import OpenAI from "openai"; import { Sandbox } from "withruntime"; const llm = new OpenAI({ baseURL: process.env.LLM_BASE_URL, apiKey: process.env.LLM_API_KEY }); const MODEL = process.env.LLM_MODEL ?? "deepseek-flash"; const REPO = "/workspace/repo"; const str = { type: "string" }; const tool = (name: string, description: string, properties: Record) => ({ type: "function" as const, function: { name, description, parameters: { type: "object", properties, required: Object.keys(properties) }, }, }); const tools = [ tool( "run", "Run a bash command in the repo. Returns the exit code and the last 6,000 characters.", { command: str }, ), tool("read_file", "Read a file. Paths are relative to the repo.", { path: str }), tool( "edit_file", "Replace one exact occurrence of `old` with `new`. An empty `old` writes a new file.", { path: str, old: str, new: str, }, ), tool("finish", "Call when the tests pass, with a short summary of the change.", { summary: str }), ]; async function runTool(sbx: Sandbox, name: string, raw: string): Promise { let a: Record; try { a = JSON.parse(raw); } catch { return `error: your arguments were not valid JSON: ${raw.slice(0, 200)}`; } const path = `${REPO}/${a.path ?? ""}`; if (name === "run") { const r = await sbx.exec(a.command ?? "", { cwd: REPO, timeoutMs: 300_000 }); return `exit ${r.exitCode}${r.timedOut ? " (timed out)" : ""}\n${(r.stdout + r.stderr).slice(-6000)}`; } if (name === "read_file") return (await sbx.files.exists(path)) ? (await sbx.files.readText(path)).slice(0, 20_000) : "error: no such file"; if (name === "edit_file") { if (!a.old) return (await sbx.files.write(path, a.new ?? ""), "written"); const text = await sbx.files.readText(path); const count = text.split(a.old).length - 1; if (count !== 1) return `error: \`old\` occurs ${count} times; copy it exactly, with more lines around it`; await sbx.files.write( path, text.replace(a.old, () => a.new ?? ""), ); return "edited"; } return `error: there is no tool called ${name}`; } export async function solve(sbx: Sandbox, task: string, maxTurns = 40): Promise { const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [ { role: "system", content: `You fix bugs in the repository at ${REPO}. Run its tests first, then read, edit and rerun until they pass, then call finish. Never edit the tests.`, }, { role: "user", content: task }, ]; for (let turn = 0; turn < maxTurns; turn++) { const reply = await llm.chat.completions.create({ model: MODEL, messages, tools }); const message = reply.choices[0]!.message; messages.push(message); // the whole message, reasoning included if (!message.tool_calls?.length) { messages.push({ role: "user", content: "Continue with a tool call, or call finish." }); continue; } for (const call of message.tool_calls) { if (call.type !== "function") continue; if (call.function.name === "finish") return call.function.arguments; const content = await runTool(sbx, call.function.name, call.function.arguments); messages.push({ role: "tool", tool_call_id: call.id, content }); } } return `stopped after ${maxTurns} turns`; } ``` ```python check import json import os from openai import OpenAI from withruntime import Sandbox llm = OpenAI(base_url=os.environ["LLM_BASE_URL"], api_key=os.environ["LLM_API_KEY"]) MODEL = os.environ.get("LLM_MODEL", "deepseek-flash") REPO = "/workspace/repo" def tool(name: str, description: str, *params: str) -> dict: props = {p: {"type": "string"} for p in params} return {"type": "function", "function": {"name": name, "description": description, "parameters": {"type": "object", "properties": props, "required": list(params)}}} TOOLS = [ tool("run", "Run a bash command in the repo. Returns the exit code and the last 6,000 characters.", "command"), tool("read_file", "Read a file. Paths are relative to the repo.", "path"), tool("edit_file", "Replace one exact occurrence of `old` with `new`. An empty `old` writes a new file.", "path", "old", "new"), tool("finish", "Call when the tests pass, with a short summary of the change.", "summary"), ] def run_tool(sbx, name: str, raw: str) -> str: try: a = json.loads(raw) except json.JSONDecodeError: return f"error: your arguments were not valid JSON: {raw[:200]}" path = f"{REPO}/{a.get('path', '')}" if name == "run": r = sbx.exec(a.get("command", ""), cwd=REPO, timeout_ms=300_000) return f"exit {r.exit_code}{' (timed out)' if r.timed_out else ''}\n{(r.stdout + r.stderr)[-6000:]}" if name == "read_file": return sbx.files.read_text(path)[:20_000] if sbx.files.exists(path) else "error: no such file" if name == "edit_file": if not a.get("old"): sbx.files.write(path, a.get("new", "")) return "written" text = sbx.files.read_text(path) if (count := text.count(a["old"])) != 1: return f"error: `old` occurs {count} times; copy it exactly, with more lines around it" sbx.files.write(path, text.replace(a["old"], a.get("new", ""), 1)) return "edited" return f"error: there is no tool called {name}" def solve(sbx, task: str, max_turns: int = 40) -> str: messages = [ {"role": "system", "content": f"You fix bugs in the repository at {REPO}. Run its tests first, then read, " "edit and rerun until they pass, then call finish. Never edit the tests."}, {"role": "user", "content": task}, ] for _ in range(max_turns): message = llm.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS).choices[0].message messages.append(message.model_dump(exclude_none=True)) # the whole message, reasoning included if not message.tool_calls: messages.append({"role": "user", "content": "Continue with a tool call, or call finish."}) continue for call in message.tool_calls: if call.function.name == "finish": return call.function.arguments content = run_tool(sbx, call.function.name, call.function.arguments) messages.append({"role": "tool", "tool_call_id": call.id, "content": content}) return f"stopped after {max_turns} turns" ``` ## Which lines make an open model reliable? The ones that turn a model's mistake into a message instead of a crash. Smaller and open models make the same mistakes as large ones, just more often, and a loop that throws on the first one never gets to see the model recover: - **Bad JSON goes back as an error.** Now and then a model emits arguments with an unescaped quote or a trailing comma. Returning the parse error, with the start of what it sent, gets a corrected call on the next turn. - **A reply with no tool call gets a nudge, not an exit.** Some models "think out loud" in plain text for a turn. Treating that as the end of the task stops the agent halfway; one short user message gets it moving again. - **Output is cut from the end.** Test runners print the failure last, so the last 6,000 characters carry the useful part. Cutting the head also keeps a long build log from filling a smaller context window. `exec` itself keeps at most 64 KiB of each stream. - **The replacement is a function.** `text.replace(old, () => new)` in TypeScript stops `$&` or `$1` in the model's code from being read as a replacement pattern, a bug that corrupts files only when the code contains a dollar sign. - **Every call has a timeout.** A command that waits for input or starts a server in the foreground returns after five minutes with `timed out`, and the model learns to background it. If your provider returns tool calls inside the message text rather than in `tool_calls`, the fix is on the serving side: vLLM, for example, needs a tool parser for the model family switched on when the server starts. ## How do you know the agent actually fixed it? Run the tests yourself, before and after, and check that the tests were not the thing that changed. `finish` is the model's opinion. A model under pressure to make a failing test pass will sometimes edit the test, so the grade below refuses any change under `tests/`: ```ts check import { Sandbox } from "withruntime"; const REPO = "/workspace/repo"; const TESTS = "python3 -m pytest -q -x"; await using sbx = await Sandbox.create({ timeoutSeconds: 3600, labels: { task: "fix-dates" } }); await sbx.exec(`git clone --depth 1 https://github.com/your-org/your-repo ${REPO}`, { check: true, timeoutMs: 300_000, }); await sbx.exec("pip install -q -e '.[test]'", { cwd: REPO, check: true, timeoutMs: 600_000 }); // Before: the task is real only if the tests fail now. const before = await sbx.exec(TESTS, { cwd: REPO, timeoutMs: 600_000 }); // ... await solve(sbx, "tests/test_dates.py fails on leap years. Fix the code.") runs here ... // After: your code decides, not the model. const after = await sbx.exec(TESTS, { cwd: REPO, timeoutMs: 600_000 }); const touchedTests = await sbx.exec("git diff --name-only -- tests", { cwd: REPO }); const diff = await sbx.exec("git diff", { cwd: REPO }); console.log({ failedBefore: before.exitCode !== 0, passesAfter: after.exitCode === 0, testsUntouched: touchedTests.stdout.trim() === "", patch: diff.stdout.slice(0, 2000), }); ``` ```python check from withruntime import Sandbox REPO, TESTS = "/workspace/repo", "python3 -m pytest -q -x" with Sandbox.create(timeout_seconds=3600, labels={"task": "fix-dates"}) as sbx: sbx.exec(f"git clone --depth 1 https://github.com/your-org/your-repo {REPO}", check=True, timeout_ms=300_000) sbx.exec("pip install -q -e '.[test]'", cwd=REPO, check=True, timeout_ms=600_000) before = sbx.exec(TESTS, cwd=REPO, timeout_ms=600_000) # the task is real only if this fails # ... solve(sbx, "tests/test_dates.py fails on leap years. Fix the code.") runs here ... after = sbx.exec(TESTS, cwd=REPO, timeout_ms=600_000) # your code decides, not the model touched_tests = sbx.exec("git diff --name-only -- tests", cwd=REPO).stdout.strip() patch = sbx.exec("git diff", cwd=REPO).stdout[:2000] print({"failed_before": before.exit_code != 0, "passes_after": after.exit_code == 0, "tests_untouched": touched_tests == "", "patch": patch}) ``` A result counts only when all three are true: the tests failed before, pass after, and the test files are untouched. Run the same task with two models and you have a fair comparison, since both worked on identical machines ([agent evals](/use-cases/agent-evals-and-swe-bench)). ## How do you choose a model for your own code? Test candidates on your own bugs, not on a public leaderboard. Your git history already holds the test set: 1. Pick ten recent commits that fixed a bug and added or changed a test. 2. For each, check out the commit's parent, then copy in only the commit's test files. The tests now fail for the reason the bug report gave. 3. Give each model the commit message's first line as the task, run `solve`, and grade with the three checks above. 4. Record pass or fail, turns used and tokens spent per model. Run every model and task in its own sandbox at the same time, so ten tasks on three models take as long as the slowest one, not thirty in a row. A paid account runs 100 sandboxes at once. Ten tasks is small, but it is enough to show whether a cheaper open model handles the kind of bug your team actually writes, which is the only question that matters for the bill. ## Where should the model's API key live? On your machine, not in the sandbox. The loop above runs on your server and only tool calls cross into the sandbox, so nothing the model writes can read `LLM_API_KEY`. If you later move the loop inside the sandbox to run it unattended, store the key as a Runtime secret bound to the provider's host: the sandbox sees a placeholder that works only on requests to that host ([secrets](/docs/security#secrets-sandboxes-never-see)). ## What does a task cost on the sandbox side? Very little next to the tokens, because the sandbox mostly waits. Take a task that keeps a 2 vCPU, 4 GiB sandbox for 15 minutes, with 90 CPU-seconds of installs and test runs: ``` CPU: 90 s / 3,600 × $0.025 = $0.0006 Memory: 15 min / 60 × 4 GiB × $0.0075 = $0.0075 Total per task: $0.0081 ``` A thousand such tasks come to $8.13 of sandbox time. CPU is billed on what the commands use, not on the vCPUs the sandbox holds, so the minutes the model spends generating cost only memory ([pricing](/pricing)). ## In short - The agent is a loop over an OpenAI-compatible API; DeepSeek, Qwen, OpenRouter and your own vLLM server differ only in three environment variables. - Four tools are enough: `run`, `read_file`, an exact-match `edit_file`, and `finish`. - Reliability on open models comes from turning mistakes into messages: bad JSON, missing tool calls and huge outputs all go back to the model. - Grade with your own test run before and after, and refuse changes to the tests. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Every task above runs in its own Firecracker microVM, ready 102 ms after the create request on Runtime's servers, with the repo, its packages and its processes kept from turn to turn. [Start with the quick start](/docs/start) and point the loop at the open model you already pay for.