AI agents that run for hours: time limits, checkpoints and recovery
An agent that runs for hours needs a time limit on every machine, a checkpoint after every step, and a supervisor that restarts it.
On Runtime (withruntime.com) an agent started with spawn belongs to its sandbox, not to your connection, so your server can deploy or crash mid-run, and a replacement sandbox for a failed run is running 102 ms after the request on Runtime's servers.
Short agent tasks fail rarely enough to retry by hand. A six-hour run meets
every failure there is: a deploy restarts your worker, the model API answers
529 for ten minutes, the agent loops on one test, the context window fills.
This post shows how to make an hours-long run survive all of that, with a
crash you can reproduce, the supervisor that recovers from it, and what the
safety net costs.
What goes wrong in a run that lasts hours?
Everything that goes wrong in a short run, multiplied by the hours, plus two failures that only long runs have: a full context window and a time limit.
| Failure | How you notice | What you do | What is lost |
|---|---|---|---|
| Your worker restarts | Nothing; the agent keeps running | Reconnect by the ids you stored | Nothing |
| Network drop to the API | A call errors | Retry; the agent never noticed | Nothing |
| Model API overloaded | The agent's own retries | Back off inside the agent, minutes not seconds | Time |
| Agent stuck in a loop | Checkpoint unchanged for 15 minutes | Kill it, restart from the checkpoint with a note | Work since the checkpoint |
| Agent crashes | Process exited with an error | Restart from the checkpoint | Work since the checkpoint |
| Context window full | The agent's own token count | Summarize into the checkpoint, start a fresh context | Nothing, if the summary is good |
| Machine time limit reached | The sandbox is no longer running | New sandbox from the checkpoint, if budget is left | Work since the checkpoint |
| The machine fails | The sandbox does not answer | New sandbox from the checkpoint | Work since the checkpoint |
Read the last column. Every failure costs at most the work since the last checkpoint, as long as two things hold: the checkpoint is written after every step, and a copy of it lives off the machine. Everything else in this post follows from that.
Why does a time limit help a long run?
Because a run with no ceiling has no way to fail cheaply. An agent stuck in a
loop at 2 a.m. keeps paying until someone wakes up. Set timeoutSeconds on
every sandbox, and the worst case has a number on it. When
the limit arrives the machine's run is over, whatever the agent was doing, so
treat it exactly like a crash: the supervisor starts a fresh sandbox from the
last checkpoint if the run's overall budget allows.
That separates two limits that are easy to mix up:
- The machine limit is
timeoutSecondson each sandbox: how long one machine may run before it is replaced. - The run budget lives in your supervisor: total hours, total restarts and total spend for the task, across every machine it uses.
A daily spending limit on the account sits above both, as the ceiling for a bad day across every run.
What goes in a checkpoint?
Everything the agent would need to continue in a new process with an empty memory: the plan, the index of the next step, a short summary of what it has learned, and the keys of every outside action it has taken. Code changes go to a git branch, committed and pushed at each step. The checkpoint file stays small, a few kilobytes, so it is cheap to copy off the machine every minute.
Two rules make checkpoints safe:
- Write it atomically. Write a temporary file and rename it over the old one. A crash in the middle of a write then leaves the previous checkpoint, never half of a new one.
- Make outside actions idempotent. A crash can land between doing something and recording it. Give each step's outside action a key that is fixed for that step, so the retry finds the action already done.
The second rule is the one people skip. This program runs as written and shows why it matters: the agent opens a pull request in step 4, dies before saving, and the retry would open a second one without the key:
TypeScriptimport { existsSync, mkdtempSync, readFileSync, renameSync, writeFileSync } from "node:fs";import { tmpdir } from "node:os";import { join } from "node:path";const STATE = join(mkdtempSync(join(tmpdir(), "agent-")), "checkpoint.json");const STEPS = [ "read the code", "write the migration", "run the tests", "open the pull request", "post a summary",];type Checkpoint = { next: number; notes: string[] };// The outside world: an API that takes idempotency keys, or a check such as// "does this branch already have a pull request?"const opened = new Set<string>();function act(step: number, key: string): string { if (step < 3) return "done"; if (opened.has(key)) return `skipped, already done under ${key}`; opened.add(key); return `done under ${key}`;}function load(): Checkpoint { return existsSync(STATE) ? JSON.parse(readFileSync(STATE, "utf8")) : { next: 0, notes: [] };}function save(cp: Checkpoint) { writeFileSync(`${STATE}.tmp`, JSON.stringify(cp)); renameSync(`${STATE}.tmp`, STATE); // atomic: a crash mid-write keeps the old checkpoint}function run(attempt: number, dieAfter?: number) { const cp = load(); console.log(`attempt ${attempt}: resuming at step ${cp.next + 1}`); for (let step = cp.next; step < STEPS.length; step++) { const result = act(step, `run-981-step-${step + 1}`); // the key is fixed per step, not per attempt console.log(` step ${step + 1}, ${STEPS[step]}: ${result}`); if (step === dieAfter) return console.log(" the process dies before saving the checkpoint"); cp.next = step + 1; cp.notes.push(STEPS[step]!); save(cp); } console.log(` finished; ${opened.size} outside actions taken in total`);}run(1, 3);run(2);Pythonimport jsonimport osimport tempfileSTATE = os.path.join(tempfile.mkdtemp(prefix="agent-"), "checkpoint.json")STEPS = ["read the code", "write the migration", "run the tests", "open the pull request", "post a summary"]# The outside world: an API that takes idempotency keys, or a check such as# "does this branch already have a pull request?"opened = set()def act(step, key): if step < 3: return "done" if key in opened: return f"skipped, already done under {key}" opened.add(key) return f"done under {key}"def load(): if os.path.exists(STATE): with open(STATE) as f: return json.load(f) return {"next": 0, "notes": []}def save(cp): with open(STATE + ".tmp", "w") as f: json.dump(cp, f) os.replace(STATE + ".tmp", STATE) # atomic: a crash mid-write keeps the old checkpointdef run(attempt, die_after=None): cp = load() print(f"attempt {attempt}: resuming at step {cp['next'] + 1}") for step in range(cp["next"], len(STEPS)): result = act(step, f"run-981-step-{step + 1}") # the key is fixed per step, not per attempt print(f" step {step + 1}, {STEPS[step]}: {result}") if step == die_after: print(" the process dies before saving the checkpoint") return cp["next"] = step + 1 cp["notes"].append(STEPS[step]) save(cp) print(f" finished; {len(opened)} outside actions taken in total")run(1, 3)run(2)It prints:
Textattempt 1: resuming at step 1 step 1, read the code: done step 2, write the migration: done step 3, run the tests: done step 4, open the pull request: done under run-981-step-4 the process dies before saving the checkpointattempt 2: resuming at step 4 step 4, open the pull request: skipped, already done under run-981-step-4 step 5, post a summary: done under run-981-step-5 finished; 2 outside actions taken in totalThe checkpoint said step 4 was still to do, which was true from its point of
view. The key is what stopped a duplicate. GitHub has no idempotency keys for
pull requests, so there the check is a lookup: does this branch already have
an open pull request? Stripe and many other APIs take an Idempotency-Key
header directly.
A checkpoint also cures a full context window. When the agent's history nears its limit, it writes a summary into the checkpoint and exits with a code that tells the supervisor to restart it. The new process reads the plan and the summary instead of four hours of tool output, and usually works better for it.
What does the supervisor look like?
A loop in your backend that copies the checkpoint off the machine every minute, watches for progress, and replaces the sandbox when the agent dies, stalls or runs out of time:
TypeScriptimport { Sandbox } from "withruntime";const RUN = "run-981";const REPO = "https://github.com/acme/app.git";const AGENT = "python3 /workspace/agent/main.py --resume /workspace/checkpoint.json";const MACHINE_SECONDS = 3600; // each sandbox's time limit: an hourconst MAX_RESTARTS = 5; // the run's budgetconst STALL_MS = 15 * 60_000;let saved = ""; // load the last copy from your database hereasync function start(checkpoint: string) { const sbx = await Sandbox.create({ vcpu: 2, memoryMiB: 4096, timeoutSeconds: MACHINE_SECONDS, labels: { run: RUN }, }); const branch = `agent/${RUN}`; await sbx.exec( `git clone ${REPO} app && cd app && (git switch ${branch} || git switch -c ${branch})`, { cwd: "/workspace", check: true, }, ); await sbx.files.upload("./agent", "/workspace/agent"); if (checkpoint) await sbx.files.write("/workspace/checkpoint.json", checkpoint); const agent = await sbx.spawn(AGENT, { cwd: "/workspace/app" }); return { sbx, agentId: agent.id }; // store both, so a restarted supervisor reconnects}let { sbx, agentId } = await start(saved);let lastProgress = Date.now();for (let restarts = 0; ;) { await new Promise((resolve) => setTimeout(resolve, 60_000)); try { const checkpoint = await sbx.files.readText("/workspace/checkpoint.json"); if (checkpoint !== saved) { saved = checkpoint; // and write it to your database: the copy off the machine lastProgress = Date.now(); } const agent = await sbx.processes.get(agentId); if (agent.state === "exited" && agent.exitCode === 0) break; // finished if (agent.state !== "running") throw new Error(`agent ${agent.state}, exit ${agent.exitCode}`); if (Date.now() - lastProgress > STALL_MS) throw new Error("no progress for 15 minutes"); } catch (error) { if (++restarts > MAX_RESTARTS) throw error; // out of budget: a person looks console.log(`restart ${restarts} from the last checkpoint: ${error}`); await sbx.pause().catch(() => {}); // paused, its files stay for a look later ({ sbx, agentId } = await start(saved)); lastProgress = Date.now(); }}await sbx.stop();Pythonimport timefrom withruntime import SandboxRUN = "run-981"REPO = "https://github.com/acme/app.git"AGENT = "python3 /workspace/agent/main.py --resume /workspace/checkpoint.json"MACHINE_SECONDS = 3600 # each sandbox's time limit: an hourMAX_RESTARTS = 5 # the run's budgetSTALL_SECONDS = 15 * 60saved = "" # load the last copy from your database heredef start(checkpoint): sbx = Sandbox.create(vcpu=2, memory_mib=4096, timeout_seconds=MACHINE_SECONDS, labels={"run": RUN}) branch = f"agent/{RUN}" sbx.exec(f"git clone {REPO} app && cd app && (git switch {branch} || git switch -c {branch})", cwd="/workspace", check=True) sbx.files.upload("./agent", "/workspace/agent") if checkpoint: sbx.files.write("/workspace/checkpoint.json", checkpoint) agent = sbx.spawn(AGENT, cwd="/workspace/app") return sbx, agent.id # store both, so a restarted supervisor reconnectssbx, agent_id = start(saved)last_progress = time.monotonic()restarts = 0while True: time.sleep(60) try: checkpoint = sbx.files.read_text("/workspace/checkpoint.json") if checkpoint != saved: saved = checkpoint # and write it to your database: the copy off the machine last_progress = time.monotonic() agent = sbx.process(agent_id) if agent.state == "exited" and agent.exit_code == 0: break # finished if agent.state != "running": raise RuntimeError(f"agent {agent.state}, exit {agent.exit_code}") if time.monotonic() - last_progress > STALL_SECONDS: raise RuntimeError("no progress for 15 minutes") except Exception as error: restarts += 1 if restarts > MAX_RESTARTS: raise # out of budget: a person looks print(f"restart {restarts} from the last checkpoint: {error}") try: sbx.pause() # paused, its files stay for a look later except Exception: pass sbx, agent_id = start(saved) last_progress = time.monotonic()sbx.stop()One catch covers every failure in the table, because they all look the
same from here: the machine stopped answering, the agent stopped running, or
the checkpoint stopped changing. A paused sandbox keeps its files, so you
can open the one that stalled and read its logs after the run is safe. The
long-running agents guide covers
reattaching to the agent's output stream from another process.
When should you snapshot instead?
When the state that matters is not in files: a database the agent filled, a dev server it configured, a warm cache it took an hour to build. A snapshot keeps the whole machine, memory and running processes included, and a new sandbox starts from it in 440 ms on Runtime's servers. Taking one pauses the sandbox for the capture, about 2.54 s for a fresh one, so take it between steps, not during a model call. For most coding agents a git branch plus a small checkpoint file is enough and costs nothing to keep.
What does the safety net cost?
Almost nothing next to the run. Take a six-hour run on 2 vCPU and 4 GiB with the agent's tools averaging 0.3 of a vCPU, and two restarts that each spend three minutes cloning and reinstalling:
TextThe run: 6 h × (0.3 vCPU × $0.025 + 4 GiB × $0.0075) = $0.2250Two restarts: 2 × 0.05 h × (2 vCPU × $0.025 + 4 GiB × $0.0075) = $0.0080The restarts add about 4% to the machine cost. The model's tokens for six hours will cost far more than either line; what the checkpoint really saves is the hours of model work a restart from zero would repeat (pricing).
In short
- Every failure in a long run costs at most the work since the last checkpoint, if the checkpoint is written each step and copied off the machine.
- Put a time limit on every sandbox and keep the run's own budget, hours, restarts and spend, in your supervisor.
- Write checkpoints atomically and give every outside action a key fixed per step, so a retry never repeats it.
- A supervisor that copies the checkpoint every minute and restarts on death, stall or time limit handles all of it with one code path.
- Snapshot when the state lives in memory or processes; a git branch and a small file cover most coding agents.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers
for an agent that mostly waits on a model (compare costs).
Runtime runs an agent started with spawn independently of your connection,
lets you give every sandbox a time limit, and starts a replacement
in 102 ms on Runtime's servers. Snapshots bring a whole machine
back in 440 ms. Start with 100 hours of a 2 vCPU, 4 GB sandbox included every month, no card:
sign in or read get started.