Reproducible agent evals: why every task needs a fresh machine
An agent eval is reproducible when each attempt starts on an identical, recorded machine and you run enough trials to see past the noise.
On Runtime (withruntime.com) a fresh sandbox from a saved snapshot is running 440 ms after the request on Runtime's servers, so giving every attempt its own clean machine costs less than a second per task. That matters because the cheapest shortcut in an eval harness, reusing a machine between tasks, is also the one that quietly changes scores. This post lists where eval noise comes from, shows how a reused machine makes a score depend on task order, and gives you the statistics to tell a real improvement from luck.
Why do agent eval scores move between runs?
Because an agent run touches far more of the world than a unit test does: a file system, a network, a clock and a sampled model. Each one is a source of variance, and only one of them is supposed to be in the score.
| Source of noise | What it does to the score | How to hold it still |
|---|---|---|
| Leftover state | Task 7 passes because task 6 installed a package or left a file | A new machine for every attempt |
| Package drift | pip install gets a newer version next week and a test breaks |
Install into the image once; pin the image version |
| Live internet | The agent finds the public fix on GitHub instead of writing it | Internet off, or an allow list you record |
| Shared CPU | A busy neighbor slows tests past their timeout | A fixed size, generous timeouts, reserved CPU for timing |
| Time | Code that reads today's date behaves differently next month | Record the date; do not depend on the clock |
| Harness drift | A prompt or tool change lands between the runs you compare | Record the harness commit in the run manifest |
| Model sampling | The same agent solves a task on one try and not the next | Several trials per task, and the right statistics |
The first six are engineering problems with clean fixes. The seventh is not a bug at all: it is the thing you are measuring, and it needs statistics, not isolation.
How does a reused machine change a score?
It makes each task's result depend on the tasks before it. Agents are messy
tenants. In a typical run an agent installs packages globally, writes to
~/.cache, edits ~/.gitconfig, starts a dev server on port 8000 and
leaves it running, and drops scratch files in /tmp.
Run three tasks on one machine and watch what leaks:
- Task A asks for a fix that needs
requests. The agent runspip install requestsand passes. - Task B needs
requeststoo, but its environment forgot to list it. On a fresh machine the agent would have to notice and install it. Here it is already there, so B passes for a reason that has nothing to do with the agent. - Task C starts its own server on port 8000. A's server still holds the port, the bind fails, and C fails for a reason that also has nothing to do with the agent.
Now shuffle the order. B before A fails, C before A passes, and the total score moves by two tasks out of three. You can test any harness this way: run the suite twice on reused machines, once forward and once reversed, and compare per-task results. A difference is contamination. On a fresh machine per attempt the order cannot matter, because nothing survives between tasks.
Cleaning up between tasks instead, with git clean and a reset, misses the
home folder, /tmp, global packages, running processes and anything the
agent wrote outside the repository. A new microVM from
the same snapshot is the only reset that is complete by construction.
What does a fresh machine mean, exactly?
The same bytes, the same size and the same network, recorded so a later run
can prove it matched. On Runtime that is four values: a pinned image version
or snapshot id, a vCPU and memory size, the network rules and a time limit.
Build the task environment once as an image, name@version pins one build
(custom images), and start every attempt from it.
Hold the network still too. With the internet on, the agent can read the answer to a public benchmark task from the repository that fixed it, and a package registry can serve a different version than last week. Turn it off for scoring, or allow only the hosts the task needs and record the list (restrict agent internet access). If the agent calls a model API, allow only that host.
What goes in the run manifest?
Everything you would need to run the same eval again and get the same machines: one record per run, plus one line per attempt. Here is a harness that writes both and gives every attempt its own sandbox:
TypeScriptimport { writeFile } from "node:fs/promises";import { Runtime } from "withruntime";type Task = { id: string; prompt: string; testCommand: string };const tasks: Task[] = JSON.parse(process.env.TASKS_JSON ?? "[]");const TRIALS = 3;const runtime = new Runtime({ waitForCapacityMs: 600_000 }); // past the limit, wait for roomconst manifest = { run: `eval-${new Date().toISOString().slice(0, 10)}`, image: "task-env@7", // a pinned build, never a moving name size: { vcpu: 2, memoryMiB: 4096 }, network: { internet: true, allow: ["api.anthropic.com"] }, // the model API only agent: { command: process.env.AGENT_COMMAND!, model: "claude-opus-5-5" }, harness: process.env.GIT_COMMIT ?? "unknown", startedAt: new Date().toISOString(),};async function attempt(task: Task, trial: number) { await using sbx = await runtime.sandboxes.create({ image: manifest.image, ...manifest.size, network: manifest.network, timeoutSeconds: 1800, labels: { run: manifest.run, task: task.id, trial: String(trial) }, }); await sbx.files.write("/workspace/TASK.md", task.prompt); const agent = await sbx.exec(manifest.agent.command, { cwd: "/workspace/repo", timeoutMs: 1_200_000, }); const tests = await sbx.exec(task.testCommand, { cwd: "/workspace/repo", timeoutMs: 300_000 }); return { task: task.id, trial, sandbox: sbx.id, pass: tests.exitCode === 0, agentTimedOut: agent.timedOut, };}const jobs = tasks.flatMap((task) => Array.from({ length: TRIALS }, (_, trial) => attempt(task, trial)),);const results = await Promise.all(jobs);await writeFile(`${manifest.run}.json`, JSON.stringify({ manifest, results }, null, 2));console.log(results.filter((r) => r.pass).length, "of", results.length, "attempts passed");Pythonimport jsonimport osfrom concurrent.futures import ThreadPoolExecutorfrom datetime import date, datetime, timezonefrom withruntime import Runtimetasks = json.loads(os.environ.get("TASKS_JSON", "[]")) # id, prompt, test_commandTRIALS = 3runtime = Runtime(wait_for_capacity=600) # past the limit, wait for roommanifest = { "run": f"eval-{date.today().isoformat()}", "image": "task-env@7", # a pinned build, never a moving name "size": {"vcpu": 2, "memory_mib": 4096}, "network": {"internet": True, "allow": ["api.anthropic.com"]}, # the model API only "agent": {"command": os.environ["AGENT_COMMAND"], "model": "claude-opus-5-5"}, "harness": os.environ.get("GIT_COMMIT", "unknown"), "started_at": datetime.now(timezone.utc).isoformat(),}def attempt(job): task, trial = job with runtime.sandboxes.create( image=manifest["image"], **manifest["size"], network=manifest["network"], timeout_seconds=1800, labels={"run": manifest["run"], "task": task["id"], "trial": str(trial)}, ) as sbx: sbx.files.write("/workspace/TASK.md", task["prompt"]) agent = sbx.exec(manifest["agent"]["command"], cwd="/workspace/repo", timeout_ms=1_200_000) tests = sbx.exec(task["test_command"], cwd="/workspace/repo", timeout_ms=300_000) return {"task": task["id"], "trial": trial, "sandbox": sbx.id, "pass": tests.exit_code == 0, "agent_timed_out": agent.timed_out}jobs = [(task, trial) for task in tasks for trial in range(TRIALS)]with ThreadPoolExecutor(max_workers=50) as pool: results = list(pool.map(attempt, jobs))with open(f"{manifest['run']}.json", "w") as f: json.dump({"manifest": manifest, "results": results}, f, indent=2)print(sum(r["pass"] for r in results), "of", len(results), "attempts passed")The sandbox id in each result line is worth keeping. Labels tie every
sandbox to its run, task and trial, so
runtime.sandboxes.list({ labels: { run } }) finds them again while they
exist (labels), so you can open the
machine behind a surprising result. await using and with end each sandbox when its attempt returns, even
when a test throws.
How do you find flaky tasks before they flake your score?
Run the known-good solution several times per task, and the empty patch once.
A task whose reference fix does not pass five times out of five is measuring
the environment, not the agent; fix it or drop it, and say so in the
manifest. A task whose tests pass with no change at all is broken the other
way. Both runs are cheap next to a full agent run, because no model is
involved, and with fresh machines their results are as trustworthy as the
real thing. On Harbor, the oracle agent runs each
task's reference solution this way, without calling a model.
How many trials do you need to trust a difference?
More than most teams run, and the comparison has to be paired by task. With one trial per task, the 95% interval on a pass rate near 50% is about ±10 points at 100 tasks and ±4.4 points at 500. A two-point gain on 500 tasks is well inside that.
Two mistakes make it worse. Treating every trial as independent makes the interval too narrow, because trials of one task are correlated: the hard tasks are hard every time. And comparing two agents' separate intervals throws away the fact that they ran the same tasks. This program, which runs as written, shows both:
TypeScript// Three trials per task, "1" for a pass. Same 20 tasks, same order, two agents.const A = "111 110 000 111 101 111 000 011 111 100 111 111 001 000 111 110 111 010 111 000".split( " ",);const B = "111 100 000 111 001 110 000 001 111 000 111 101 000 000 111 100 110 000 111 000".split( " ",);const mean = (xs: number[]) => xs.reduce((s, x) => s + x, 0) / xs.length;const sd = (xs: number[]) => { const m = mean(xs); return Math.sqrt(xs.reduce((s, x) => s + (x - m) ** 2, 0) / (xs.length - 1));};const rate = (trials: string) => [...trials].filter((c) => c === "1").length / trials.length;const pct = (x: number) => `${(100 * x).toFixed(1)}%`;function report(name: string, runs: string[]) { const perTask = runs.map(rate); const p = mean(perTask); const trials = runs.join("").length; const naive = 1.96 * Math.sqrt((p * (1 - p)) / trials); // treats every trial as independent const honest = (1.96 * sd(perTask)) / Math.sqrt(perTask.length); // tasks are the unit console.log(`${name}: ${pct(p)} naive ±${pct(naive)} per-task ±${pct(honest)}`);}report("Agent A", A);report("Agent B", B);const diff = A.map((a, i) => rate(a) - rate(B[i]!));const d = mean(diff);const half = (1.96 * sd(diff)) / Math.sqrt(diff.length);const verdict = d - half > 0 ? "A is better" : "not distinguishable yet";console.log(`A - B, paired by task: ${pct(d)} ±${pct(half)}: ${verdict}`);Pythonfrom statistics import mean, stdev# Three trials per task, "1" for a pass. Same 20 tasks, same order, two agents.A = "111 110 000 111 101 111 000 011 111 100 111 111 001 000 111 110 111 010 111 000".split()B = "111 100 000 111 001 110 000 001 111 000 111 101 000 000 111 100 110 000 111 000".split()def rate(trials): return trials.count("1") / len(trials)def pct(x): return f"{100 * x:.1f}%"def report(name, runs): per_task = [rate(r) for r in runs] p = mean(per_task) trials = len("".join(runs)) naive = 1.96 * (p * (1 - p) / trials) ** 0.5 # treats every trial as independent honest = 1.96 * stdev(per_task) / len(per_task) ** 0.5 # tasks are the unit print(f"{name}: {pct(p)} naive ±{pct(naive)} per-task ±{pct(honest)}")report("Agent A", A)report("Agent B", B)diff = [rate(a) - rate(b) for a, b in zip(A, B)]d = mean(diff)half = 1.96 * stdev(diff) / len(diff) ** 0.5verdict = "A is better" if d - half > 0 else "not distinguishable yet"print(f"A - B, paired by task: {pct(d)} ±{pct(half)}: {verdict}")It prints:
TextAgent A: 63.3% naive ±12.2% per-task ±17.7%Agent B: 46.7% naive ±12.6% per-task ±18.6%A - B, paired by task: 16.7% ±7.5%: A is betterLook at the separate intervals and the agents overlap; you would call it a tie. Pair them by task and the difference is clear, because B never does better than A on any task. The naive interval is about a third narrower than the honest one, which is how teams end up shipping a change that was noise. This only works if every trial ran on the same kind of machine; leftover state would add a correlation you cannot remove afterward.
What does a reproducible run cost?
Take 500 tasks, three trials each, ten minutes per attempt on 2 vCPU and 4 GiB, with tests and tools using 120 CPU-seconds, plus five reference runs per task at two minutes and 60 CPU-seconds each:
TextAgent CPU: 1,500 × 120 s / 3,600 × $0.025 = $1.25Agent memory: 1,500 × 600 s / 3,600 × 4 GiB × $0.0075 = $7.50Reference CPU: 2,500 × 60 s / 3,600 × $0.025 = $1.04Reference mem: 2,500 × 120 s / 3,600 × 4 GiB × $0.0075 = $2.50Total: $12.29The model's tokens cost far more than the machines, which is the point: once fresh machines cost this little, there is no reason left to share them. The agent evals use case has the SWE-bench loop itself, and Harbor runs Terminal-Bench tasks on Runtime unchanged.
In short
- Seven things move agent eval scores; six are fixed by running each attempt on a fresh, recorded machine.
- A reused machine makes results depend on task order: run the suite forward and reversed to catch it.
- Record the image version, size, network rules, model and harness commit with every run.
- Run the reference fix five times per task to find flaky tasks before they reach your score.
- Compare agents paired by task, with tasks as the unit, not trials.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). A Runtime sandbox starts from a snapshot in 440 ms on Runtime's servers, and a paid account runs 100 at once, with creates past that waiting for room. Each one is a Firecracker microVM of its own, billed on the CPU your tests use (pricing). Start with 100 free hours, no card: sign in or read get started.