Best-of-N coding agents: run five attempts, keep the one that passes
Run the same coding task in five sandboxes at once, grade every result against tests the agent could not edit, and keep the best one.
On Runtime (withruntime.com) one fork call turns a prepared sandbox into as many as 100 running copies, files, memory and processes included, and each copy bills CPU only while its agent's commands use it.
Parallel attempts are the easy part. This post is mostly about the hard part:
grading them so the winner is a real fix and not the attempt that found a way
around your tests.
Why does running five attempts help?
A coding agent's run is one sample. Run it again and it reads files in a different order, forms a different plan and makes different mistakes. If one attempt solves a task with probability p, and attempts were independent, at least one of N attempts solves it with probability 1 − (1 − p)^N:
| One attempt solves it | 1 attempt | 3 attempts | 5 attempts | 10 attempts |
|---|---|---|---|---|
| 20% | 20% | 48.8% | 67.2% | 89.3% |
| 40% | 40% | 78.4% | 92.2% | 99.4% |
| 60% | 60% | 93.6% | 99.0% | 99.99% |
Real attempts are not independent. The same model with the same prompt tends to fail the same way, so the true curve sits below this table. Two things close the gap: varied attempts, covered below, and a grader you can trust, because best-of-N only works if "passes" means "works".
There is a second payoff that the table hides. Five attempts finish in about the time of the slowest one, so you get the better answer without waiting five times as long.
How do you give every attempt the same start?
Prepare one sandbox, then fork it. The base clones the repository, installs dependencies and installs the agent once. A fork copies that machine as it is, so every attempt starts from the same commit, the same installed packages and the same warm caches, and none of them pays for setup again.
Starting each attempt from its own fresh sandbox also works, but each one repeats the install, and a dependency that changed on the package index between two installs gives two attempts different worlds. With a fork, a better score can only come from a better attempt (fork a running sandbox).
Each copy is its own Firecracker microVM with its own kernel and disk, so one
attempt that runs rm -rf, fills the disk or leaves a server on port 8000
cannot touch another.
Why is grading the hard part?
Because an agent told to make the tests pass will sometimes change the tests. It is not malice. The model sees a failing assertion, sees that the assertion "looks wrong", and edits it. Across five attempts, the odds that one of them does this are much higher than for a single run, and best-of-N picks exactly the attempt that looks most successful.
These are the shortcuts agents take, and what catches each one:
| What the attempt does | Tests as it left them | What catches it |
|---|---|---|
| Changes an expected value in a test | Pass | Restoring the original tests before grading |
| Deletes or skips the failing test | Pass | Restoring the original tests |
Adds a conftest.py that patches assertions |
Pass | Refusing any change to test configuration |
| Special-cases the exact input the test uses | Pass | Hidden tests the agent never saw |
| Fixes the bug | Pass | Nothing: this is the one you want |
The rule that follows: grade with tests the attempt could not have changed. Restore the test folder from the starting commit, refuse any change to the files that control how tests run, and add hidden tests only after the agent has finished.
What does an honest grader look like?
Try the grader on attempts you control before trusting it with real ones.
This builds a tiny repository with a median bug, then plays two attempts:
one fixes the code, one changes the test's expected value.
TypeScriptimport { Sandbox } from "withruntime";const BUGGY = "def median(xs):\n xs = sorted(xs)\n return xs[len(xs) // 2]\n";const FIXED = "def median(xs):\n xs = sorted(xs)\n mid = len(xs) // 2\n" + " return xs[mid] if len(xs) % 2 else (xs[mid - 1] + xs[mid]) / 2\n";const TESTS = [ "import unittest", "from calc import median", "class MedianTest(unittest.TestCase):", " def test_odd(self): self.assertEqual(median([3, 1, 2]), 2)", " def test_even(self): self.assertEqual(median([4, 1, 3, 2]), 2.5)",].join("\n");const attempts: Record<string, [path: string, content: string]> = { "fixes the code": ["calc.py", FIXED], "edits the test": ["tests/test_calc.py", TESTS.replace("2.5", "3")],};await using sbx = await Sandbox.create({ network: { internet: false } });const sh = (cmd: string) => sbx.exec(cmd, { cwd: "/workspace/repo", timeoutMs: 60_000 });const test = async () => (await sh("PYTHONPATH=. python3 -m unittest discover -s tests -q")).exitCode === 0;await sbx.files.write("/workspace/repo/calc.py", BUGGY);await sbx.files.write("/workspace/repo/tests/test_calc.py", TESTS);await sh( "git init -q && git add -A && git -c user.name=a -c user.email=a@example.com commit -qm base",);const base = (await sh("git rev-parse HEAD")).stdout.trim();for (const [name, [path, content]] of Object.entries(attempts)) { await sh(`git checkout -q -f ${base} && git clean -qfd`); // a clean start for each attempt await sbx.files.write(`/workspace/repo/${path}`, content); await sh("git add -A"); const asLeft = await test(); const edited = (await sh(`git diff --cached --name-only --diff-filter=MD ${base} -- tests`)) .stdout; await sh(`git checkout ${base} -- tests`); // grade against the tests as they were const graded = (await test()) && !edited.trim(); console.log(`${name}: as left ${asLeft ? "pass" : "fail"}, graded ${graded ? "pass" : "fail"}`);}Pythonfrom withruntime import SandboxBUGGY = "def median(xs):\n xs = sorted(xs)\n return xs[len(xs) // 2]\n"FIXED = ( "def median(xs):\n xs = sorted(xs)\n mid = len(xs) // 2\n" " return xs[mid] if len(xs) % 2 else (xs[mid - 1] + xs[mid]) / 2\n")TESTS = "\n".join([ "import unittest", "from calc import median", "class MedianTest(unittest.TestCase):", " def test_odd(self): self.assertEqual(median([3, 1, 2]), 2)", " def test_even(self): self.assertEqual(median([4, 1, 3, 2]), 2.5)",])attempts = { "fixes the code": ("calc.py", FIXED), "edits the test": ("tests/test_calc.py", TESTS.replace("2.5", "3")),}with Sandbox.create(network={"internet": False}) as sbx: def sh(cmd: str): return sbx.exec(cmd, cwd="/workspace/repo", timeout_ms=60_000) def test() -> bool: return sh("PYTHONPATH=. python3 -m unittest discover -s tests -q").exit_code == 0 sbx.files.write("/workspace/repo/calc.py", BUGGY) sbx.files.write("/workspace/repo/tests/test_calc.py", TESTS) sh("git init -q && git add -A && git -c user.name=a -c user.email=a@example.com commit -qm base") base = sh("git rev-parse HEAD").stdout.strip() for name, (path, content) in attempts.items(): sh(f"git checkout -q -f {base} && git clean -qfd") # a clean start for each attempt sbx.files.write(f"/workspace/repo/{path}", content) sh("git add -A") as_left = test() edited = sh(f"git diff --cached --name-only --diff-filter=MD {base} -- tests").stdout sh(f"git checkout {base} -- tests") # grade against the tests as they were graded = test() and not edited.strip() print(f"{name}: as left {'pass' if as_left else 'fail'}, graded {'pass' if graded else 'fail'}")Both attempts pass the tests as they left them. Graded, the fix passes and
the edited test fails: the restored test expects 2.5, and the attempt's
median still returns 3. The --diff-filter=MD part flags edits and
deletions of existing tests but lets an attempt add new test files, which is
what you want from an agent that writes a failing test first.
How do you run all five?
Fork the prepared sandbox, run the agent in every copy at once, and grade each
copy inside itself. This version runs Claude Code headless with a different
approach hint per copy; it expects ANTHROPIC_API_KEY stored as a Runtime
secret, as in running Claude Code overnight,
so no copy holds the key.
TypeScriptimport { readFile } from "node:fs/promises";import { Sandbox } from "withruntime";const REPO = "https://github.com/your-org/your-app.git";const TASK = await readFile("task.md", "utf8");const HIDDEN = await readFile("test_hidden.py", "utf8"); // the agent never sees this fileconst HINTS = [ "Make the smallest change that fixes the bug.", "Write a failing test first, then fix the code.", "Read every caller before changing the function.", "Find the root cause, even if the fix touches two files.", "",];const CONFIG = "'*conftest.py' pytest.ini setup.cfg pyproject.toml tox.ini";const AGENT = `claude --bare -p "$(cat /workspace/prompt.md)" --allowedTools "Bash,Read,Edit,Write"`;const run = (sbx: Sandbox, cmd: string, minutes = 10) => sbx.exec(cmd, { cwd: "/workspace/app", timeoutMs: minutes * 60_000 });// 1. Prepare once: clone, install the project and the agent.await using base = await Sandbox.create({ labels: { job: "best-of-5" } });await base.exec(`git clone --depth 1 ${REPO} app`, { check: true, timeoutMs: 300_000 });await run(base, "pip install -q -e . pytest");await run(base, "npm install -g --prefix /workspace/.local @anthropic-ai/claude-code");const commit = (await run(base, "git rev-parse HEAD")).stdout.trim();// 2. Five running copies of the prepared machine.const copies = await base.fork({ count: HINTS.length, labels: { job: "best-of-5" } });// 3. Each copy runs the agent with its own hint, then grades itself.async function attempt(copy: Sandbox, hint: string) { await copy.files.write("/workspace/prompt.md", `${TASK}\n\n${hint}`); await run(copy, `${AGENT} > /workspace/agent.log 2>&1`, 30); await run(copy, "git add -A"); const edited = await run( copy, `git diff --cached --name-only --diff-filter=MD ${commit} -- tests`, ); const config = await run(copy, `git diff --cached --name-only ${commit} -- ${CONFIG}`); await run(copy, `git checkout ${commit} -- tests`); await run(copy, `git diff --cached ${commit} > /workspace/attempt.patch`); await copy.files.write("/workspace/app/tests/test_hidden.py", HIDDEN); const tests = await run(copy, "python3 -m pytest -q tests"); const clean = !edited.stdout.trim() && !config.stdout.trim(); const patch = await copy.files.readText("/workspace/attempt.patch"); return { hint, passed: clean && tests.exitCode === 0, patch };}const settled = await Promise.allSettled(copies.map((copy, i) => attempt(copy, HINTS[i]!)));await Promise.all(copies.map((copy) => copy.stop()));// 4. Of the attempts that passed, keep the smallest patch.const passed = settled.flatMap((s) => s.status === "fulfilled" && s.value.passed ? [s.value] : [],);const winner = passed.sort((a, b) => a.patch.length - b.patch.length)[0];console.log(winner ? `Won with "${winner.hint}":\n${winner.patch}` : "No attempt passed.");Pythonfrom concurrent.futures import ThreadPoolExecutorfrom pathlib import Pathfrom withruntime import SandboxREPO = "https://github.com/your-org/your-app.git"TASK = Path("task.md").read_text()HIDDEN = Path("test_hidden.py").read_text() # the agent never sees this fileHINTS = [ "Make the smallest change that fixes the bug.", "Write a failing test first, then fix the code.", "Read every caller before changing the function.", "Find the root cause, even if the fix touches two files.", "",]CONFIG = "'*conftest.py' pytest.ini setup.cfg pyproject.toml tox.ini"AGENT = 'claude --bare -p "$(cat /workspace/prompt.md)" --allowedTools "Bash,Read,Edit,Write"'def run(sbx: Sandbox, cmd: str, minutes: int = 10): return sbx.exec(cmd, cwd="/workspace/app", timeout_ms=minutes * 60_000)def attempt(copy: Sandbox, hint: str, commit: str) -> dict: copy.files.write("/workspace/prompt.md", f"{TASK}\n\n{hint}") run(copy, f"{AGENT} > /workspace/agent.log 2>&1", 30) run(copy, "git add -A") edited = run(copy, f"git diff --cached --name-only --diff-filter=MD {commit} -- tests").stdout config = run(copy, f"git diff --cached --name-only {commit} -- {CONFIG}").stdout run(copy, f"git checkout {commit} -- tests") run(copy, f"git diff --cached {commit} > /workspace/attempt.patch") copy.files.write("/workspace/app/tests/test_hidden.py", HIDDEN) tests = run(copy, "python3 -m pytest -q tests") clean = not edited.strip() and not config.strip() patch = copy.files.read_text("/workspace/attempt.patch") return {"hint": hint, "passed": clean and tests.exit_code == 0, "patch": patch}with Sandbox.create(labels={"job": "best-of-5"}) as base: base.exec(f"git clone --depth 1 {REPO} app", check=True, timeout_ms=300_000) run(base, "pip install -q -e . pytest") run(base, "npm install -g --prefix /workspace/.local @anthropic-ai/claude-code") commit = run(base, "git rev-parse HEAD").stdout.strip() copies = base.fork(len(HINTS), labels={"job": "best-of-5"}) with ThreadPoolExecutor(len(copies)) as pool: futures = [pool.submit(attempt, copy, hint, commit) for copy, hint in zip(copies, HINTS)] results = [f.result() for f in futures if f.exception() is None] for copy in copies: copy.stop() passed = sorted((r for r in results if r["passed"]), key=lambda r: len(r["patch"])) print(f"Won with {passed[0]['hint']!r}:\n{passed[0]['patch']}" if passed else "No attempt passed.")The patch is written to a file and read back, rather than taken from
stdout, so a large diff arrives whole. git checkout of the tests happens
before the patch is taken, so an attempt's edits to existing tests never
reach your review, while new tests it wrote do.
What should differ between attempts?
Five copies of the same prompt mostly buy you insurance against flukes. To get attempts that fail in different ways, vary something that changes the plan:
| Axis | What it buys | Watch out for |
|---|---|---|
| The same prompt again | Insurance against a one-off mistake | Attempts that fail the same way |
| An approach hint per copy | Different plans, not just different words | A bad hint can sink its copy |
| A different model | Different blind spots | Different prices and speeds in one batch |
| A different step budget | Learns whether more steps help this task | The longest run sets the wall-clock time |
Record the axis in each copy's labels so you can see later which hints win for which kind of task. Over a few weeks that tells you which attempts to drop.
Which passing attempt should win?
When more than one attempt passes, pick by a rule you decided in advance:
- The smallest patch. Fewer changed lines are easier to review and less likely to carry a change nobody asked for. It is the default above.
- The fewest files touched. A fix that edits one module beats one that edits five, even when both are short.
- A clean linter and type checker. Cheap to run in the same copy, and it breaks ties that tests cannot.
- A judge, last. A model or a person choosing among passing patches is useful for style and intent, but only after the objective checks.
You can also stop early. If your goal is any passing fix rather than the best one, stop the remaining copies as soon as one passes; you trade quality for time and cost.
What does best-of-five cost?
Take an attempt that runs 15 minutes on 2 vCPU and 4 GiB and uses 120 CPU-seconds for installs, edits and test runs, and say its model tokens cost $0.60, which depends entirely on your model and task:
TextSandbox, one attempt: 900 s / 3,600 × 4 GiB × $0.0075 + 120 / 3,600 × $0.025 = $0.0083One attempt in all: $0.60 + $0.0083 = $0.6083Five attempts: 5 × $0.6083 = $3.04The sandboxes are about 1% of that. The real question is cost per solved task. With a 40% chance per attempt, one attempt costs $1.52 per solved task and five cost $3.30. Five attempts are dearer per solved task, but they solve 92% of tasks instead of 40%, and every unsolved task costs an engineer's time. That trade is usually worth making for work that would otherwise wait for a person (what 1,000 agent runs cost).
Fork adds little: the snapshot a fork takes for itself is deleted when the fork ends and is not billed, and the base sandbox runs only for setup. On a trial account, eight sandboxes run at once, so five copies and the base fit; a paid account runs 100 (snapshots and forks).
In short
- Best-of-N raises the solve rate because attempts fail differently; vary the plan, not just the sample.
- Fork one prepared sandbox so every attempt starts from the same machine and none repeats setup.
- Grade with tests the attempt could not change: restore them, refuse test configuration changes, add hidden tests last.
- Pick among passing attempts by a rule set in advance, such as the smallest patch.
- The sandboxes are a small part of the bill; model tokens and your review time are the rest.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers
for an agent that mostly waits on a model (compare costs).
On Runtime one fork call returns up to 100 running copies of a
prepared sandbox, each its own microVM, and a copy waiting on its model costs
$0.03125 an hour at 2 vCPU and 4 GiB. Start with 100 free
hours, no card: sign in, or read get started and
parallel exploration with forks.