Cold starts in agent loops: why milliseconds per tool call add up
An agent makes dozens of tool calls per task, so per-call overhead multiplies; keep one warm sandbox per task and run the loop near it.
On Runtime (withruntime.com) a command in a running sandbox takes 53 ms at the median on Runtime's servers, and a new sandbox runs its first command 221 ms after the create request. A single cold start rarely matters. Forty of them in a row, each with a setup step and a network round trip, can matter more than the model. This post works through the arithmetic for one task, shows where the time really goes, and gives a small harness to measure your own loop.
Where does an agent's time go?
Into three places: the model thinking, the tools running, and the overhead of getting each tool call to a machine and back. The first two are the work. The third is what this post is about, and it scales with the number of calls, not with how hard the task is.
Take one ordinary coding task: 40 tool calls, one after another, each a short
command such as grep, cat or a single test file. Say the model takes 4
seconds per turn, which is typical for a mid-sized model answering with a tool
call. The model's share is fixed:
TextModel time: 40 turns × 4 s = 160 sEverything below is added on top of those 160 seconds, and the user waits for all of it.
What does each architecture add?
It depends on how often the loop starts a machine, and how much setup each start repeats. Here are four common shapes, with Runtime's median server-side times:
A new sandbox per tool call, with setup. Some designs treat each call as a function: create, clone, install, run, delete. With 20 seconds of setup:
Text40 × (221 ms + 20 s) = 808.8 sA new sandbox per tool call, no setup. Pure computation with nothing to install, such as a calculator tool:
Text40 × 221 ms = 8.8 sOne sandbox per task. Create once, then send every call to the same running machine:
Text221 ms + 40 × 53 ms = 2.3 sOne sandbox per conversation, paused between messages. Say the 40 calls come in 8 user messages. Each message wakes the sandbox once, and the other calls go to a running machine:
Text8 × 153 ms + 32 × 53 ms = 2.9 sSet against the 160 seconds the model takes:
| Architecture | Overhead added | Share of the model's time |
|---|---|---|
| New sandbox per call, with setup | 808.8 s | 506% |
| New sandbox per call, no setup | 8.8 s | 5.5% |
| One sandbox per task | 2.3 s | 1.5% |
| One per conversation, paused between | 2.9 s | 1.8% |
The lesson is not the cold start. A fast start keeps even the per-call design to a few percent. What dominates is the setup that each fresh machine repeats, because a fresh machine has forgotten everything the last call did. Keep state, and the overhead almost disappears.
Why does the network count 40 times?
Because every tool call is a round trip from your agent loop to the sandbox and back, and the loop waits for each one. The figures above are measured on Runtime's servers. Your loop adds the distance between it and them, once per call.
Runtime runs in one region, in Virginia. From a laptop in the US Mountain time zone, the same command took 121 ms instead of 53 ms. Over 40 calls:
TextExtra per task: 40 × (121 ms − 53 ms) = 2.7 sFrom another continent, add your own round-trip time to that, once per call. The fix is cheap: run the agent loop on a server in a US East region, close to the sandboxes, and let only the final answer travel to the user. A model API call is one round trip per turn either way, so moving the loop costs nothing there.
Why does the slow tail matter more than the median?
Because a task with many calls almost always hits a slow one. If each call has a 5% chance of landing at or beyond the 95th percentile, a task of 40 calls has at least one such call with probability:
Text1 − 0.95^40 = 87%So most of your tasks feel the p95, and long sessions feel the p99. That is why a provider's median alone tells you little about an agent loop (what p95 latency means). Ask for the tail, and measure it. On Runtime's servers a command in a running sandbox took 156 ms at p95 and 196 ms at p99, and a new sandbox's first command 484 ms at p95 (speed).
How do you measure your own loop?
Time every tool call where your loop makes it, then look at the distribution, not the average. This harness runs a sequence of commands the way an agent would, one after another, and reports the median, the p95 and the total:
TypeScriptimport { Sandbox } from "withruntime";const calls = [ "ls", "cat /etc/os-release", "python3 -c 'print(sum(range(1000)))'", "grep -c processor /proc/cpuinfo", "df -h /workspace",];await using sbx = await Sandbox.create();const times: number[] = [];for (let round = 0; round < 8; round++) { for (const command of calls) { const started = performance.now(); await sbx.exec(command, { timeoutMs: 30_000 }); times.push(performance.now() - started); }}times.sort((a, b) => a - b);const at = (q: number) => Math.round(times[Math.min(times.length - 1, Math.floor(q * times.length))]!);const total = Math.round(times.reduce((sum, t) => sum + t, 0));console.log(`${times.length} calls: median ${at(0.5)} ms, p95 ${at(0.95)} ms, total ${total} ms`);Pythonimport timefrom withruntime import Sandboxcalls = [ "ls", "cat /etc/os-release", "python3 -c 'print(sum(range(1000)))'", "grep -c processor /proc/cpuinfo", "df -h /workspace",]with Sandbox.create() as sbx: times: list[float] = [] for _ in range(8): for command in calls: started = time.perf_counter() sbx.exec(command, timeout_ms=30_000) times.append((time.perf_counter() - started) * 1000) times.sort() def at(q: float) -> int: return round(times[min(len(times) - 1, int(q * len(times)))]) print(f"{len(times)} calls: median {at(0.5)} ms, p95 {at(0.95)} ms, total {round(sum(times))} ms")Run it from where your agent loop runs in production, not from your laptop, and compare the total with your model's time for the same number of turns. If the sandbox share is more than a few percent, one of the fixes below will find it.
What actually shrinks the overhead?
Most of the gain comes from the first two items; the rest are refinements.
- Keep one sandbox per task or conversation. Never one per call. State survives, setup runs once, and each call costs a command, not a start.
- Bake setup into an image. Dependencies installed at build time cost nothing at create time, so even the first call of a task is fast (custom images).
- Run the loop near the sandboxes. One round trip per call, multiplied by every call, is the overhead you control most cheaply.
- Pause between messages, not between calls. A wake costs more than a command, so pause when the user is reading, and let the 60 seconds idle pause catch the rest (pause, don't delete).
- Run parallel tool calls in parallel. When a model asks for three reads at once, send them concurrently. The overhead becomes the slowest call, not the sum.
- Let one call do more. A tool that accepts a short script instead of a
single command turns five round trips into one. Models handle
cd app && npm test -- --run authas easily as three separate calls.
Item 6 has a limit: a call that does too much returns too much output, and the model reads all of it. Combine steps that always run together; keep steps whose output the model needs to decide the next one separate.
In short
- Per-call overhead multiplies by the number of tool calls, so a task of 40 calls feels it 40 times.
- The cold start is small; repeated setup on fresh machines is what costs minutes.
- Every call is a round trip, so run the agent loop near the sandboxes.
- Many calls means most tasks hit the p95: measure the tail, not just the median.
- Keep one sandbox per task or conversation, bake setup into an image, and pause only between messages.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). On Runtime's servers a command in a running sandbox takes 53 ms, a new sandbox runs its first command 221 ms after the create request, and a paused one runs its next command 153 ms after the wake request. CPU bills at $0.025 per vCPU-hour of actual use, so a sandbox kept for a whole task costs little while the model thinks. Start with 100 free hours, no card: sign in, or read get started and the pricing.