Runtime

AI agents that run for hours: time limits, checkpoints and recovery

An agent that runs for hours needs a time limit on every machine, a checkpoint after every step, and a supervisor that restarts it.

On Runtime (withruntime.com) an agent started with spawn belongs to its sandbox, not to your connection, so your server can deploy or crash mid-run, and a replacement sandbox for a failed run is running 102 ms after the request on Runtime's servers. Short agent tasks fail rarely enough to retry by hand. A six-hour run meets every failure there is: a deploy restarts your worker, the model API answers 529 for ten minutes, the agent loops on one test, the context window fills. This post shows how to make an hours-long run survive all of that, with a crash you can reproduce, the supervisor that recovers from it, and what the safety net costs.

What goes wrong in a run that lasts hours?

Everything that goes wrong in a short run, multiplied by the hours, plus two failures that only long runs have: a full context window and a time limit.

Failure How you notice What you do What is lost
Your worker restarts Nothing; the agent keeps running Reconnect by the ids you stored Nothing
Network drop to the API A call errors Retry; the agent never noticed Nothing
Model API overloaded The agent's own retries Back off inside the agent, minutes not seconds Time
Agent stuck in a loop Checkpoint unchanged for 15 minutes Kill it, restart from the checkpoint with a note Work since the checkpoint
Agent crashes Process exited with an error Restart from the checkpoint Work since the checkpoint
Context window full The agent's own token count Summarize into the checkpoint, start a fresh context Nothing, if the summary is good
Machine time limit reached The sandbox is no longer running New sandbox from the checkpoint, if budget is left Work since the checkpoint
The machine fails The sandbox does not answer New sandbox from the checkpoint Work since the checkpoint

Read the last column. Every failure costs at most the work since the last checkpoint, as long as two things hold: the checkpoint is written after every step, and a copy of it lives off the machine. Everything else in this post follows from that.

Why does a time limit help a long run?

Because a run with no ceiling has no way to fail cheaply. An agent stuck in a loop at 2 a.m. keeps paying until someone wakes up. Set timeoutSeconds on every sandbox, and the worst case has a number on it. When the limit arrives the machine's run is over, whatever the agent was doing, so treat it exactly like a crash: the supervisor starts a fresh sandbox from the last checkpoint if the run's overall budget allows.

That separates two limits that are easy to mix up:

  • The machine limit is timeoutSeconds on each sandbox: how long one machine may run before it is replaced.
  • The run budget lives in your supervisor: total hours, total restarts and total spend for the task, across every machine it uses.

A daily spending limit on the account sits above both, as the ceiling for a bad day across every run.

What goes in a checkpoint?

Everything the agent would need to continue in a new process with an empty memory: the plan, the index of the next step, a short summary of what it has learned, and the keys of every outside action it has taken. Code changes go to a git branch, committed and pushed at each step. The checkpoint file stays small, a few kilobytes, so it is cheap to copy off the machine every minute.

Two rules make checkpoints safe:

  1. Write it atomically. Write a temporary file and rename it over the old one. A crash in the middle of a write then leaves the previous checkpoint, never half of a new one.
  2. Make outside actions idempotent. A crash can land between doing something and recording it. Give each step's outside action a key that is fixed for that step, so the retry finds the action already done.

The second rule is the one people skip. This program runs as written and shows why it matters: the agent opens a pull request in step 4, dies before saving, and the retry would open a second one without the key:

TypeScriptimport { existsSync, mkdtempSync, readFileSync, renameSync, writeFileSync } from "node:fs";import { tmpdir } from "node:os";import { join } from "node:path";const STATE = join(mkdtempSync(join(tmpdir(), "agent-")), "checkpoint.json");const STEPS = [  "read the code",  "write the migration",  "run the tests",  "open the pull request",  "post a summary",];type Checkpoint = { next: number; notes: string[] };// The outside world: an API that takes idempotency keys, or a check such as// "does this branch already have a pull request?"const opened = new Set<string>();function act(step: number, key: string): string {  if (step < 3) return "done";  if (opened.has(key)) return `skipped, already done under ${key}`;  opened.add(key);  return `done under ${key}`;}function load(): Checkpoint {  return existsSync(STATE) ? JSON.parse(readFileSync(STATE, "utf8")) : { next: 0, notes: [] };}function save(cp: Checkpoint) {  writeFileSync(`${STATE}.tmp`, JSON.stringify(cp));  renameSync(`${STATE}.tmp`, STATE); // atomic: a crash mid-write keeps the old checkpoint}function run(attempt: number, dieAfter?: number) {  const cp = load();  console.log(`attempt ${attempt}: resuming at step ${cp.next + 1}`);  for (let step = cp.next; step < STEPS.length; step++) {    const result = act(step, `run-981-step-${step + 1}`); // the key is fixed per step, not per attempt    console.log(`  step ${step + 1}, ${STEPS[step]}: ${result}`);    if (step === dieAfter) return console.log("  the process dies before saving the checkpoint");    cp.next = step + 1;    cp.notes.push(STEPS[step]!);    save(cp);  }  console.log(`  finished; ${opened.size} outside actions taken in total`);}run(1, 3);run(2);
Pythonimport jsonimport osimport tempfileSTATE = os.path.join(tempfile.mkdtemp(prefix="agent-"), "checkpoint.json")STEPS = ["read the code", "write the migration", "run the tests", "open the pull request", "post a summary"]# The outside world: an API that takes idempotency keys, or a check such as# "does this branch already have a pull request?"opened = set()def act(step, key):    if step < 3:        return "done"    if key in opened:        return f"skipped, already done under {key}"    opened.add(key)    return f"done under {key}"def load():    if os.path.exists(STATE):        with open(STATE) as f:            return json.load(f)    return {"next": 0, "notes": []}def save(cp):    with open(STATE + ".tmp", "w") as f:        json.dump(cp, f)    os.replace(STATE + ".tmp", STATE)  # atomic: a crash mid-write keeps the old checkpointdef run(attempt, die_after=None):    cp = load()    print(f"attempt {attempt}: resuming at step {cp['next'] + 1}")    for step in range(cp["next"], len(STEPS)):        result = act(step, f"run-981-step-{step + 1}")  # the key is fixed per step, not per attempt        print(f"  step {step + 1}, {STEPS[step]}: {result}")        if step == die_after:            print("  the process dies before saving the checkpoint")            return        cp["next"] = step + 1        cp["notes"].append(STEPS[step])        save(cp)    print(f"  finished; {len(opened)} outside actions taken in total")run(1, 3)run(2)

It prints:

Textattempt 1: resuming at step 1  step 1, read the code: done  step 2, write the migration: done  step 3, run the tests: done  step 4, open the pull request: done under run-981-step-4  the process dies before saving the checkpointattempt 2: resuming at step 4  step 4, open the pull request: skipped, already done under run-981-step-4  step 5, post a summary: done under run-981-step-5  finished; 2 outside actions taken in total

The checkpoint said step 4 was still to do, which was true from its point of view. The key is what stopped a duplicate. GitHub has no idempotency keys for pull requests, so there the check is a lookup: does this branch already have an open pull request? Stripe and many other APIs take an Idempotency-Key header directly.

A checkpoint also cures a full context window. When the agent's history nears its limit, it writes a summary into the checkpoint and exits with a code that tells the supervisor to restart it. The new process reads the plan and the summary instead of four hours of tool output, and usually works better for it.

What does the supervisor look like?

A loop in your backend that copies the checkpoint off the machine every minute, watches for progress, and replaces the sandbox when the agent dies, stalls or runs out of time:

TypeScriptimport { Sandbox } from "withruntime";const RUN = "run-981";const REPO = "https://github.com/acme/app.git";const AGENT = "python3 /workspace/agent/main.py --resume /workspace/checkpoint.json";const MACHINE_SECONDS = 3600; // each sandbox's time limit: an hourconst MAX_RESTARTS = 5; // the run's budgetconst STALL_MS = 15 * 60_000;let saved = ""; // load the last copy from your database hereasync function start(checkpoint: string) {  const sbx = await Sandbox.create({    vcpu: 2,    memoryMiB: 4096,    timeoutSeconds: MACHINE_SECONDS,    labels: { run: RUN },  });  const branch = `agent/${RUN}`;  await sbx.exec(    `git clone ${REPO} app && cd app && (git switch ${branch} || git switch -c ${branch})`,    {      cwd: "/workspace",      check: true,    },  );  await sbx.files.upload("./agent", "/workspace/agent");  if (checkpoint) await sbx.files.write("/workspace/checkpoint.json", checkpoint);  const agent = await sbx.spawn(AGENT, { cwd: "/workspace/app" });  return { sbx, agentId: agent.id }; // store both, so a restarted supervisor reconnects}let { sbx, agentId } = await start(saved);let lastProgress = Date.now();for (let restarts = 0; ;) {  await new Promise((resolve) => setTimeout(resolve, 60_000));  try {    const checkpoint = await sbx.files.readText("/workspace/checkpoint.json");    if (checkpoint !== saved) {      saved = checkpoint; // and write it to your database: the copy off the machine      lastProgress = Date.now();    }    const agent = await sbx.processes.get(agentId);    if (agent.state === "exited" && agent.exitCode === 0) break; // finished    if (agent.state !== "running") throw new Error(`agent ${agent.state}, exit ${agent.exitCode}`);    if (Date.now() - lastProgress > STALL_MS) throw new Error("no progress for 15 minutes");  } catch (error) {    if (++restarts > MAX_RESTARTS) throw error; // out of budget: a person looks    console.log(`restart ${restarts} from the last checkpoint: ${error}`);    await sbx.pause().catch(() => {}); // paused, its files stay for a look later    ({ sbx, agentId } = await start(saved));    lastProgress = Date.now();  }}await sbx.stop();
Pythonimport timefrom withruntime import SandboxRUN = "run-981"REPO = "https://github.com/acme/app.git"AGENT = "python3 /workspace/agent/main.py --resume /workspace/checkpoint.json"MACHINE_SECONDS = 3600  # each sandbox's time limit: an hourMAX_RESTARTS = 5  # the run's budgetSTALL_SECONDS = 15 * 60saved = ""  # load the last copy from your database heredef start(checkpoint):    sbx = Sandbox.create(vcpu=2, memory_mib=4096, timeout_seconds=MACHINE_SECONDS,                         labels={"run": RUN})    branch = f"agent/{RUN}"    sbx.exec(f"git clone {REPO} app && cd app && (git switch {branch} || git switch -c {branch})",             cwd="/workspace", check=True)    sbx.files.upload("./agent", "/workspace/agent")    if checkpoint:        sbx.files.write("/workspace/checkpoint.json", checkpoint)    agent = sbx.spawn(AGENT, cwd="/workspace/app")    return sbx, agent.id  # store both, so a restarted supervisor reconnectssbx, agent_id = start(saved)last_progress = time.monotonic()restarts = 0while True:    time.sleep(60)    try:        checkpoint = sbx.files.read_text("/workspace/checkpoint.json")        if checkpoint != saved:            saved = checkpoint  # and write it to your database: the copy off the machine            last_progress = time.monotonic()        agent = sbx.process(agent_id)        if agent.state == "exited" and agent.exit_code == 0:            break  # finished        if agent.state != "running":            raise RuntimeError(f"agent {agent.state}, exit {agent.exit_code}")        if time.monotonic() - last_progress > STALL_SECONDS:            raise RuntimeError("no progress for 15 minutes")    except Exception as error:        restarts += 1        if restarts > MAX_RESTARTS:            raise  # out of budget: a person looks        print(f"restart {restarts} from the last checkpoint: {error}")        try:            sbx.pause()  # paused, its files stay for a look later        except Exception:            pass        sbx, agent_id = start(saved)        last_progress = time.monotonic()sbx.stop()

One catch covers every failure in the table, because they all look the same from here: the machine stopped answering, the agent stopped running, or the checkpoint stopped changing. A paused sandbox keeps its files, so you can open the one that stalled and read its logs after the run is safe. The long-running agents guide covers reattaching to the agent's output stream from another process.

When should you snapshot instead?

When the state that matters is not in files: a database the agent filled, a dev server it configured, a warm cache it took an hour to build. A snapshot keeps the whole machine, memory and running processes included, and a new sandbox starts from it in 440 ms on Runtime's servers. Taking one pauses the sandbox for the capture, about 2.54 s for a fresh one, so take it between steps, not during a model call. For most coding agents a git branch plus a small checkpoint file is enough and costs nothing to keep.

What does the safety net cost?

Almost nothing next to the run. Take a six-hour run on 2 vCPU and 4 GiB with the agent's tools averaging 0.3 of a vCPU, and two restarts that each spend three minutes cloning and reinstalling:

TextThe run:      6 h × (0.3 vCPU × $0.025 + 4 GiB × $0.0075)       = $0.2250Two restarts: 2 × 0.05 h × (2 vCPU × $0.025 + 4 GiB × $0.0075) = $0.0080

The restarts add about 4% to the machine cost. The model's tokens for six hours will cost far more than either line; what the checkpoint really saves is the hours of model work a restart from zero would repeat (pricing).

In short

  • Every failure in a long run costs at most the work since the last checkpoint, if the checkpoint is written each step and copied off the machine.
  • Put a time limit on every sandbox and keep the run's own budget, hours, restarts and spend, in your supervisor.
  • Write checkpoints atomically and give every outside action a key fixed per step, so a retry never repeats it.
  • A supervisor that copies the checkpoint every minute and restarts on death, stall or time limit handles all of it with one code path.
  • Snapshot when the state lives in memory or processes; a git branch and a small file cover most coding agents.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Runtime runs an agent started with spawn independently of your connection, lets you give every sandbox a time limit, and starts a replacement in 102 ms on Runtime's servers. Snapshots bring a whole machine back in 440 ms. Start with 100 hours of a 2 vCPU, 4 GB sandbox included every month, no card: sign in or read get started.

100 hours of a 2 vCPU, 4 GB sandbox,included every month.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card