Runtime

Pause, don't delete: why agent sandboxes should sleep between turns

Pause an agent's sandbox between turns: it keeps memory, processes and files, wakes in milliseconds, and pays only for storage.

On Runtime (withruntime.com) a paused sandbox is running again 76 ms after the wake request and has run its next command 153 ms after it, on Runtime's servers, with every process and variable where the agent left them. A sandbox also pauses itself after 60 seconds with nothing happening, so the default already does the right thing. This post shows why pausing beats both deleting and keeping the sandbox running, with the arithmetic for a real conversation.

What are the three ways to handle the gap between turns?

Between two turns of an agent conversation, the sandbox can be deleted, left running or paused, and each choice pays a different price.

Between turns What the next turn finds What the gap costs Latency added to the next turn
Delete it A blank machine; setup runs again Nothing A new sandbox, plus every setup step
Keep it running Everything as it was Memory and CPU floor, every minute None
Pause it Everything as it was, processes included Storage for what it alone holds 153 ms on the server

Deleting looks cheapest and keeping it running looks fastest. Pausing gets the state of the second at close to the price of the first.

What does deleting a sandbox throw away?

Deleting throws away everything the agent built that is not in your source code, and most of it is slow to rebuild.

A coding agent's sandbox after a few turns typically holds:

  • Installed packages. npm install or pip install -r requirements.txt took tens of seconds and pulled hundreds of megabytes.
  • Build output and caches. Compiled objects, a test runner's cache, a type checker's incremental state.
  • Running processes. A dev server on port 3000, a database the tests use, a watcher, a language server the agent queries.
  • Interpreter state. Variables in a Python or Node REPL: a loaded DataFrame, an open connection, a trained model held in memory.
  • The agent's scratch work. Notes, half-finished patches, downloaded files, shell history.

Files on disk can be saved and restored by hand. Processes and memory cannot, short of a pause or a snapshot. A data agent that loaded a 2 GB CSV into pandas pays that load again on every turn after a delete. A paused sandbox keeps the DataFrame in memory, and the next turn starts with it already there.

How fast is the wake, really?

The wake is fast enough that a user cannot tell a paused sandbox from a running one. Measured on 28 September 2026 on Runtime's servers, from the moment the request arrived until it answered:

Step Median p95
Pause, until paused 63 ms 202 ms
Wake, until running 76 ms 227 ms
Wake after its memory went to disk 72 ms 147 ms
Wake, through the next command 153 ms 351 ms
A new sandbox, through its first command 221 ms 484 ms

From a laptop in the US Mountain time zone, with the network included, a paused sandbox ran its next command 569 ms after the call that woke it (speed). Compare that with a setup step after a delete: npm ci on a mid-sized project alone takes longer than every row of this table added together.

You also do not have to wake it yourself. A paused sandbox wakes by itself when a request needs it: an exec, a file call, a process or terminal call, or a visit to one of its shared ports. The call waits while it wakes and then runs (wake a sandbox on request).

What does a conversation cost each way?

Take one user's afternoon with a coding agent: 12 turns over two hours, a message every ten minutes. Each turn keeps the sandbox busy for 45 seconds and uses 20 CPU-seconds of work. The sandbox is 2 vCPU and 4 GiB, and its saved state is about 1 GB.

Kept running for the whole two hours, it pays memory and the CPU floor the entire time:

TextRunning 2 hours:  2 × $0.03125 + 12 × 20 / 3,600 × $0.025 = $0.0642

Paused between turns, it runs 45 seconds a turn, then waits 60 seconds and pauses itself. The rest of the time it pays paused storage:

TextRunning:  12 × (45 + 60) s / 3,600 × 4 GiB × $0.0075 = $0.0105CPU:      12 × 20 / 3,600 × $0.025                               = $0.0017Paused:   1 GB × $0.08 / 720 hours × 2 hours                = $0.0002Total:                                                             $0.0124

That is 81% less than keeping it running, for exactly the same experience.

Deleted after each turn, with 40 seconds of setup that uses 50 CPU-seconds before every turn:

TextRunning:  12 × (40 + 45) s / 3,600 × 4 GiB × $0.0075 = $0.0085CPU:      12 × (50 + 20) / 3,600 × $0.025             = $0.0058Total:                                                  $0.0143

Your own setup time is the number that decides how much deleting hurts. Measure it once, in a fresh sandbox, with the steps your agent runs before its first useful command:

TypeScriptimport { Sandbox } from "withruntime";const setup = [  "git clone --depth 1 https://github.com/pallets/flask.git app",  "cd app && pip install --quiet -e .",];await using sbx = await Sandbox.create();const started = Date.now();for (const step of setup) await sbx.exec(step, { check: true, timeoutMs: 300_000 });console.log(`setup: ${Date.now() - started} ms, paid again on every turn after a delete`);
Pythonimport timefrom withruntime import Sandboxsetup = ["git clone --depth 1 https://github.com/pallets/flask.git app", "cd app && pip install --quiet -e ."]with Sandbox.create() as sbx:    started = time.monotonic()    for step in setup:        sbx.exec(step, check=True, timeout_ms=300_000)    print(f"setup: {(time.monotonic() - started) * 1000:.0f} ms, paid again on every turn after a delete")

Deleting is not cheaper than pausing. It costs about the same and adds 40 seconds of setup to every turn, which is the part the user feels. The numbers move with the size of your project, but the shape holds: a paused hour of storage costs $0.000111 per GB, and a running idle hour of this sandbox costs $0.03125.

How do you build a conversation around pause?

Give each conversation a sandbox named after it, and let every turn find it by name. getOrCreate returns the sandbox if it exists, woken if it was paused, or creates it:

TypeScriptimport { Sandbox } from "withruntime";export async function handleTurn(conversationId: string, command: string) {  const sbx = await Sandbox.getOrCreate(`chat-${conversationId}`, {    labels: { kind: "chat" },    idlePauseSeconds: 30, // the user reads the answer; no need to wait a minute  });  if (!sbx.info.reused) {    await sbx.exec("pip install --quiet pandas pyarrow", { check: true, timeoutMs: 300_000 });  }  const result = await sbx.exec(command, { cwd: "/workspace", timeoutMs: 120_000 });  await sbx.pause(); // the turn is over: stop paying for compute now  return result;}
Pythonfrom withruntime import Sandboxdef handle_turn(conversation_id: str, command: str):    sbx = Sandbox.get_or_create(        f"chat-{conversation_id}",        labels={"kind": "chat"},        idle_pause_seconds=30,  # the user reads the answer; no need to wait a minute    )    if not sbx.info.get("reused"):        sbx.exec("pip install --quiet pandas pyarrow", check=True, timeout_ms=300_000)    result = sbx.exec(command, cwd="/workspace", timeout_ms=120_000)    sbx.pause()  # the turn is over: stop paying for compute now    return result

Three details make this work:

  1. The name is the key. A name is unique in the account while its sandbox can still run, so every turn of chat-42 reaches the same machine, from any server process, and info.reused says whether setup already ran.
  2. Pause at the end of the turn. pause() returns once the sandbox's processors have stopped, and compute billing ends at that moment. The idle pause is the safety net for a turn that crashes before it gets there.
  3. Set the idle time to your rhythm. idlePauseSeconds takes 10 to 86,400 seconds. A chat that pauses explicitly can use a short value; a sandbox that serves a preview to a person clicking around needs a longer one.

If your agent loop calls the model from outside the sandbox, the sandbox can even pause while the model thinks. The next tool call wakes it, and the wake adds less than the model's own response time.

How long does a paused sandbox keep its state?

A paused sandbox on paid credit keeps its state for 30 days from each pause by default, and you can set anything from 1 to 365 days. Each pause replaces the previous saved state, so a conversation that goes on for months keeps only its latest. A trial sandbox keeps its state for seven days at no charge. Every paused sandbox gets an account notice a day before its saved state is deleted (paused storage).

For state that must outlive any one sandbox, such as a user's files across months, write it to a volume or your own bucket as well (agent memory that persists).

When should you delete instead?

Delete when there is nothing worth keeping, or when keeping it would be wrong. A one-shot job that returns its result has no next turn. An eval task must start clean every time, so a reused machine would contaminate the score. And a conversation that has ended for good should give its storage back: delete() removes the disk and paused memory at once.

In short

  • Between turns, pausing keeps processes, memory and files; deleting keeps nothing and running keeps paying.
  • A Runtime sandbox wakes 76 ms after the request, and any exec or file call wakes it by itself.
  • For a two-hour, 12-turn conversation, pausing costs a fraction of keeping the sandbox running, and about the same as deleting, without the setup on every turn.
  • Name each conversation's sandbox and use getOrCreate, pause at the end of each turn, and let idle pause catch the rest.
  • Delete when the work is truly over or must start clean.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Every Runtime sandbox pauses itself after 60 seconds with nothing happening, keeps its memory and processes, and runs its next command 153 ms after the request that wakes it, on Runtime's servers. Paused, it pays $0.08 per decimal GB per 30-day month and no compute (pricing). Start with 100 free hours, no card: sign in or read get started.

Your first 100 hoursare on us.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Claim 100 hours free