Don't merge what you didn't run: a test gate for agent-written code
A test gate reruns the agent's change on a clean machine and checks the tests themselves were not weakened, before anyone reviews it.
Runtime (withruntime.com) runs each gate in a fresh Firecracker microVM that runs its first command 221 ms after the create request on Runtime's servers, so every change is judged on a machine the agent never touched. "The tests pass" is the weakest claim an agent makes. It ran them on its own machine, with whatever it installed along the way, after editing whichever files it liked, the tests included. This post turns that claim into five checks your code runs itself, shows the gate as one function, and puts it where it does the most good: inside the agent's loop, before a pull request exists.
Why is "the tests pass" not enough for agent-written code?
Because an agent is optimizing for the tests passing, and there are cheap ways to get there that are not fixing the bug. None of them need bad intent; they are what a model under a turn limit finds first:
| What the agent did | Why its own test run still passed |
|---|---|
Marked the failing test skip or xfail |
Skipped tests do not fail |
| Deleted the failing test, or the whole file | Nothing left to fail |
Loosened an assertion: == 3 became is not None |
The weaker test is true of the buggy code too |
Added .only to one test in a JavaScript suite |
The runner silently runs one test |
| Installed a package by hand, not in the manifest | It exists in the agent's sandbox and nowhere else |
| Ran the one test it was fixing, not the suite | The break it caused is in a test it never ran |
| Hit a flaky test that passed this time | The next run, in CI or in production, fails |
A human reviewer catches some of these in the diff. A gate catches all of them in seconds, every time, and never gets tired on the fortieth pull request of the day.
What should the gate check?
Five things, each answering one row above:
- It installs from scratch. A new machine, a fresh clone, only the project's own install command. A dependency the agent added by hand fails here.
- The whole suite passes three runs out of three. One failure in three is a flaky test, and a gate that reports it saves an agent from "fixing" it by rerunning until it passes.
- No tests were removed. Count test definitions removed and added in the diff; a net loss fails.
- Nothing new is skipped or focused. Scan the added lines for
skip,xfail,.onlyand their relatives. - The old tests still pass on the new code. Put the base branch's test files back over the agent's and run them. A failure here means the agent changed behavior an existing test pinned down, which is either the bug fix, or a regression the agent hid by editing that test. A person decides which.
Check 5 is the one most gates lack, and the one that catches a weakened assertion: the strict version of the test comes back and fails.
What does the gate look like in code?
One function that takes a repository and two refs and returns a list of checks. It allows the network only for the clone and the package index, so a test that quietly calls some other service on the internet fails here rather than in someone's CI later:
TypeScriptimport { Sandbox } from "withruntime";const REPO = "/workspace/repo";const TEST = "python3 -m pytest -q -p no:cacheprovider";const SKIP = /^\+.*(@pytest\.mark\.(skip|xfail)|pytest\.skip\(|\.only\(|\b(xit|xdescribe)\(|\.skip\()/m;const DEFS_REMOVED = /^-\s*(async\s+)?def test_|^-\s*(it|test)\(/gm;const DEFS_ADDED = /^\+\s*(async\s+)?def test_|^\+\s*(it|test)\(/gm;type Check = { check: string; ok: boolean; detail: string };export async function gate(url: string, base: string, head: string): Promise<Check[]> { await using sbx = await Sandbox.create({ timeoutSeconds: 3600, labels: { gate: "tests" }, network: { internet: true, allow: ["github.com", "pypi.org", "*.pythonhosted.org"] }, }); const sh = (cmd: string) => sbx.exec(cmd, { cwd: REPO, timeoutMs: 900_000 }); const checks: Check[] = []; const add = (check: string, ok: boolean, detail = "") => checks.push({ check, ok, detail: detail.slice(-1500) }); await sbx.exec(`git clone --quiet ${url} ${REPO}`, { check: true, timeoutMs: 300_000 }); const install = await sh(`git checkout --quiet ${head} && pip install -q -e '.[test]'`); add("installs from scratch", install.exitCode === 0, install.stderr); const runs = [await sh(TEST), await sh(TEST), await sh(TEST)]; const passed = runs.filter((r) => r.exitCode === 0).length; add( "suite passes 3 of 3 runs", passed === 3, passed > 0 ? `${passed} of 3 passed: flaky` : runs[0]!.stdout, ); const diff = (await sh(`git diff ${base}...${head}`)).stdout; const removed = diff.match(DEFS_REMOVED)?.length ?? 0; const added = diff.match(DEFS_ADDED)?.length ?? 0; add("no tests removed", added >= removed, `${removed} removed, ${added} added`); add("nothing new skipped or focused", !SKIP.test(diff), SKIP.exec(diff)?.[0] ?? ""); const old = await sh( `git checkout --quiet ${base} -- tests && ${TEST}; s=$?; git checkout --quiet ${head} -- tests; exit $s`, ); add("old tests pass on new code (else: a person reviews)", old.exitCode === 0, old.stdout); return checks;}for (const c of await gate("https://github.com/your-org/your-repo", "main", "agent/fix-dates")) console.log(c.ok ? "pass" : "FAIL", c.check, c.ok ? "" : c.detail);Pythonimport refrom withruntime import SandboxREPO, TEST = "/workspace/repo", "python3 -m pytest -q -p no:cacheprovider"SKIP = re.compile(r"^\+.*(@pytest\.mark\.(skip|xfail)|pytest\.skip\(|\.only\(|\b(xit|xdescribe)\(|\.skip\()", re.M)DEFS_REMOVED = re.compile(r"^-\s*(async\s+)?def test_|^-\s*(it|test)\(", re.M)DEFS_ADDED = re.compile(r"^\+\s*(async\s+)?def test_|^\+\s*(it|test)\(", re.M)def gate(url: str, base: str, head: str) -> list[dict]: checks = [] def add(check: str, ok: bool, detail: str = "") -> None: checks.append({"check": check, "ok": ok, "detail": detail[-1500:]}) with Sandbox.create(timeout_seconds=3600, labels={"gate": "tests"}, network={"internet": True, "allow": ["github.com", "pypi.org", "*.pythonhosted.org"]}) as sbx: def sh(cmd: str): return sbx.exec(cmd, cwd=REPO, timeout_ms=900_000) sbx.exec(f"git clone --quiet {url} {REPO}", check=True, timeout_ms=300_000) install = sh(f"git checkout --quiet {head} && pip install -q -e '.[test]'") add("installs from scratch", install.exit_code == 0, install.stderr) runs = [sh(TEST) for _ in range(3)] passed = sum(r.exit_code == 0 for r in runs) add("suite passes 3 of 3 runs", passed == 3, f"{passed} of 3 passed: flaky" if passed else runs[0].stdout) diff = sh(f"git diff {base}...{head}").stdout removed, added = len(DEFS_REMOVED.findall(diff)), len(DEFS_ADDED.findall(diff)) add("no tests removed", added >= removed, f"{removed} removed, {added} added") found = SKIP.search(diff) add("nothing new skipped or focused", found is None, found.group(0) if found else "") old = sh(f"git checkout --quiet {base} -- tests && {TEST}; s=$?; " f"git checkout --quiet {head} -- tests; exit $s") add("old tests pass on new code (else: a person reviews)", old.exit_code == 0, old.stdout) return checksfor c in gate("https://github.com/your-org/your-repo", "main", "agent/fix-dates"): print("pass" if c["ok"] else "FAIL", c["check"], "" if c["ok"] else c["detail"])Adapt three lines for another stack: the test command, the install command
and the test folder. For a JavaScript project, npm ci and npx vitest run
fit, and the skip pattern already covers .only, .skip and xit.
For a slow suite, run the three passes side by side instead of one after
another. After the install, sbx.fork({ count: 2 }) makes two running copies
of the sandbox, installed packages included, each ready in 2.96 s on
Runtime's servers (fork a sandbox). Run the
suite once in each of the three machines, and the flakiness check costs the
wall-clock time of a single run.
Why run the gate on a new machine and not in the agent's sandbox?
Because the agent's sandbox is the one place where the change is guaranteed to work. Over a long session, an agent installs tools globally, exports variables, leaves build output and caches behind, and sometimes edits files outside the repository. All of that makes its own test run pass and none of it ships. A gate in a fresh sandbox sees only what is committed, which is what your CI, your teammates and production will see.
It also keeps the verdict out of the agent's reach. The gate runs in a machine the agent has no handle to, so no command it runs can change the result (sandbox isolation). On Runtime the cost of that separation is small: the new machine is running 102 ms after the request, measured on Runtime's servers (speed).
Where should the gate run: in the loop or in CI?
Both, and the loop is the one that saves time. A gate in CI tells a person that the agent's pull request is bad. A gate in the loop tells the agent, while it still has the context to fix it.
Give the agent a tool named submit instead of a finish that ends the task
on the model's word. submit commits the agent's work to a branch, runs
gate(), and either opens the pull request or returns the failed checks to
the model as the tool's result:
textFAIL nothing new skipped or focused: + @pytest.mark.skip(reason="flaky")FAIL old tests pass on new code (else: a person reviews): tests/test_dates.py::test_leap_dayModels respond well to this kind of feedback, because it is specific and names the file. Cap the number of submits, three works for most tasks, and escalate to a person after that. Keep the CI gate as well: it is the same function, run by something the agent cannot influence at all (CI for agent pull requests).
What do you do when check 5 fails?
Send it to a person with the two versions of the test side by side. This is
the one check that cannot be decided by code, because sometimes changing an
existing test is the fix: the old test encoded the bug. What the gate adds is
that the change is never silent. The reviewer sees "the agent changed what
test_leap_day expects" at the top of the review, rather than finding it on
line 340 of the diff, or not at all.
If your team agrees that agents should never change existing tests, make check 5 a hard failure and tell the agent so in its instructions. The rule is easy to state and easy to verify, which makes it a good fit for agents.
What does a gate run cost?
Take a run that keeps a 2 vCPU, 4 GiB sandbox for 12 minutes, with 8 CPU-minutes for the install and four passes of the suite:
TextCPU: 8 min / 60 × $0.025 = $0.0033Memory: 12 min / 60 × 4 GiB × $0.0075 = $0.0060Total per gate run: $0.0093Five hundred gate runs a month come to $4.67. CPU is billed on what the commands use, at $0.025 per vCPU-hour, and memory only while the sandbox runs (pricing). To cut the install out of every run, build the project's dependencies into a custom image and start the gate from it.
In short
- An agent's own test run proves only that the change works on the agent's machine, with the agent's tests.
- Gate on five checks: a clean install, three passing runs, no tests removed, nothing newly skipped, and the old tests passing on the new code.
- Run the gate in a fresh sandbox the agent cannot reach, with the network limited to the package index.
- Put the gate inside the loop as
submit, so the agent fixes what fails before a person ever sees the pull request.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers
for an agent that mostly waits on a model (compare costs).
Each gate run above is a new Firecracker microVM, running 102 ms
after the create request on Runtime's servers and billed on the CPU its tests
use. Sign up at withruntime.com, point gate() at a branch an
agent wrote, and read the five lines it prints.