Stateful Python for agents: what survives a call, a pause and a crash
Give each agent one long-lived Python session, keep the cells that build its state on disk, and replay them when the session restarts.
Runtime (withruntime.com) keeps an agent's Python variables in memory between calls, through a pause and into each copy of a fork, and a paused sandbox has run its next command 153 ms after the wake request on Runtime's servers. A stateful session is what lets an agent work like an analyst: load a 2 GB file once, then ask forty small questions of it. It is also a source of bugs nobody planned for, because "stateful" means different things to memory, to disk and to your own server. A variable survives the next call but not a crash; a file survives both; a database connection may survive neither. This post maps what lasts through each event, then builds a session that recovers from the ones that wipe memory.
What exactly is the state of a Python session?
Four layers, and they live in different places:
| Layer | Lives in | Example |
|---|---|---|
| Variables, imports, functions | The Python process | df, model, import pandas as pd |
| Files | The sandbox's disk | /workspace/data/sales.csv, a parquet cache |
| Installed packages | The disk, then memory | pip install polars, then import polars |
| Connections to the outside | The process and the far end | A database cursor, an open HTTP session |
The model thinks of all four as "the session". Your code has to know which one a given event touches. Here is what survives what:
| Event | Variables | Files | Installed packages | Outside connections |
|---|---|---|---|---|
| The next call | Kept | Kept | Kept | Kept |
| A cell hits its time limit | Kept | Kept | Kept | Usually kept |
| A pause, then a wake | Kept | Kept | Kept | Often closed by the far end |
| The Python process dies | Lost | Kept | Kept, import again | Lost |
| A fork | Copied | Copied | Copied | Reconnect in each copy |
| A memory snapshot, started later | Kept | Kept | Kept | Reconnect |
| A disk-only snapshot | Lost | Kept | Kept, import again | Lost |
Two rows surprise people. A pause freezes the whole machine, so Python never notices it, but a database server on the other side sees a client that went silent and closes the connection after its own timeout. And when the process dies, usually from running out of memory, the disk is untouched: the CSV is still there, the packages are still installed, and only the variables are gone.
Why does a session need a recovery plan?
Because the process will die sometimes, and the model will not notice. An
agent that loaded a DataFrame twenty calls ago does not reload it before each
question. If the process restarted in between, the next cell fails with
NameError: name 'df' is not defined, and a model that does not know why will
start guessing: reinstall pandas, rewrite the cell, apologize.
The fix is to treat the cells that build state differently from the cells that ask questions. A setup cell imports, loads data or defines a function that later cells rely on. Keep those, in order, on the sandbox's disk. When the session restarts, replay them before running anything else, and tell the model what happened in one line.
The model knows which of its cells are setup cells, so let it say so: the tool
takes a setup flag next to the code.
What does a session that recovers look like?
The interpreter keeps a Python process per context and reports two things on
every run that matter here: contextStarted, which is true when this run had
to start a fresh process, and status, which is "lost" when the process
died during the cell. With those, recovery is a few lines:
TypeScriptimport { Sandbox } from "withruntime";type Cell = Awaited<ReturnType<Sandbox["interpreter"]["run"]>>;const LOG = "/workspace/.session/setup.json"; // on the sandbox's disk, not in your serverexport class PythonSession { constructor(private readonly sbx: Sandbox) {} // setup: this cell imports, loads data or defines something later cells rely on. async run(code: string, setup = false): Promise<string> { let cell = await this.cell(code); const earlier = await this.setupCells(); let note = ""; if ((cell.contextStarted || cell.status === "lost") && earlier.length > 0) { for (const step of earlier) await this.cell(step); note = `[The Python session restarted; ${earlier.length} setup cells were run again.]\n`; if (cell.contextStarted) cell = await this.cell(code); // it ran without its state: once more } if (setup && cell.status === "ok") await this.sbx.files.write(LOG, JSON.stringify([...earlier, code])); return note + describe(cell); } private cell(code: string): Promise<Cell> { return this.sbx.interpreter.run(code, { timeoutMs: 120_000 }); } private async setupCells(): Promise<string[]> { try { return JSON.parse(new TextDecoder().decode(await this.sbx.files.read(LOG))) as string[]; } catch { return []; // no setup yet } }}function describe(cell: Cell): string { const value = cell.results.find((r) => r.main)?.data["text/plain"]; const parts = [cell.stdout, cell.stderr, value === undefined ? "" : String(value)]; if (cell.error) parts.push(`${cell.error.name}: ${cell.error.value}`); if (cell.status === "timeout") parts.push("[Stopped at 120 s. Earlier variables are still defined.]"); if (cell.status === "lost") parts.push("[Python died during this cell, most often from lack of memory.]"); return parts.filter(Boolean).join("\n").slice(-8000) || "[no output]";}Pythonimport jsonfrom withruntime import SandboxLOG = "/workspace/.session/setup.json" # on the sandbox's disk, not in your serverclass PythonSession: def __init__(self, sbx: Sandbox): self.sbx = sbx def run(self, code: str, setup: bool = False) -> str: """setup: this cell imports, loads data or defines something later cells rely on.""" cell = self._cell(code) earlier = self._setup_cells() note = "" if (cell["contextStarted"] or cell["status"] == "lost") and earlier: for step in earlier: self._cell(step) note = f"[The Python session restarted; {len(earlier)} setup cells were run again.]\n" if cell["contextStarted"]: cell = self._cell(code) # it ran without its state: once more if setup and cell["status"] == "ok": self.sbx.files.write(LOG, json.dumps([*earlier, code])) return note + describe(cell) def _cell(self, code: str) -> dict: return self.sbx.interpreter.run(code, timeout_ms=120_000) def _setup_cells(self) -> list[str]: try: return json.loads(self.sbx.files.read(LOG)) except Exception: return [] # no setup yetdef describe(cell: dict) -> str: value = next((r["data"].get("text/plain") for r in cell["results"] if r["main"]), None) parts = [cell["stdout"], cell["stderr"], "" if value is None else str(value)] if cell["error"]: parts.append(f"{cell['error']['name']}: {cell['error']['value']}") if cell["status"] == "timeout": parts.append("[Stopped at 120 s. Earlier variables are still defined.]") if cell["status"] == "lost": parts.append("[Python died during this cell, most often from lack of memory.]") return "\n".join(p for p in parts if p)[-8000:] or "[no output]"A few decisions in there are deliberate:
- The log lives on the sandbox's disk. If your web server restarts, or the next message lands on a different instance, the setup log is still with the session it describes.
- Only setup cells are replayed. Replaying every cell would repeat writes,
uploads and slow model fits. A setup cell should be safe to run twice, and
"load this file into
df" always is. - A cell that killed the process is not run again. If it ran out of memory once, it will again. The model reads the note and splits the work.
- A timeout is not a restart. The process is interrupted and keeps its variables, so the note says so. Without it, models reload everything after every slow cell.
The rest of the tool, the result format and how to wire it to each model API, is in build a code interpreter tool for any LLM.
Why does a session die, and how do you make it rarer?
Almost always memory. Pandas makes copies freely: a filter, a merge or a
df.copy() can double what a DataFrame holds for a moment, and a 3 GB frame in
a 4 GiB sandbox will not survive a merge. The process is killed, the run ends
lost, and the session starts again.
Three habits keep it alive:
- Size the sandbox for the peak, not the file. Plan for three times the
largest frame. Memory is set at create with
memoryMiBand is the line you pay for while the sandbox runs (size CPU and memory). - Let the model see memory. Expose a cell that reports the process's
resident memory, and mention it in the tool description. A model that can
see 3.4 GB in use will
delan intermediate frame before the next merge. - Keep a disk copy of anything slow to rebuild. Write the cleaned frame to parquet once; the replayed setup cell then reads it in seconds instead of cleaning the raw data again.
Python# A cell the model can run at any time: the Python process's resident memory, in MiB.int(open("/proc/self/status").read().split("VmRSS:")[1].split()[0]) // 1024Where does the session live between messages?
In the sandbox, not in your server. The next message in a conversation can be handled by any of your servers, a day later, with nothing but the sandbox's id. While the user is away, the sandbox pauses itself after 60 seconds of quiet, keeps its memory, and any call wakes it. A paused sandbox has run its next command 153 ms after the wake request, so the user never sees the pause.
This version runs as two "servers" in one script to make the point. The second needs only the id:
TypeScriptimport { Sandbox } from "withruntime";// Server one: the conversation starts. The data and the setup log go on the sandbox's disk.const first = await Sandbox.create({ memoryMiB: 4096, labels: { conversation: "c-8812" } });await first.files.write("/workspace/data/sales.csv", "region,revenue\nnorth,1200\nsouth,950\n");const setup = ["import pandas as pd", "df = pd.read_csv('/workspace/data/sales.csv')"];await first.files.write("/workspace/.session/setup.json", JSON.stringify(setup));// Server two, later: the next message arrives somewhere else, with only the id.await using again = await Sandbox.connect(first.id);const log = await again.files.read("/workspace/.session/setup.json");const cells = JSON.parse(new TextDecoder().decode(log)) as string[];console.log(`${again.id}: ${cells.length} setup cells on disk, ready to replay if needed`);Pythonimport jsonfrom withruntime import Sandbox# Server one: the conversation starts. The data and the setup log go on the sandbox's disk.first = Sandbox.create(memory_mib=4096, labels={"conversation": "c-8812"})first.files.write("/workspace/data/sales.csv", "region,revenue\nnorth,1200\nsouth,950\n")setup = ["import pandas as pd", "df = pd.read_csv('/workspace/data/sales.csv')"]first.files.write("/workspace/.session/setup.json", json.dumps(setup))# Server two, later: the next message arrives somewhere else, with only the id.with Sandbox.connect(first.id) as again: cells = json.loads(again.files.read("/workspace/.session/setup.json")) print(f"{again.id}: {len(cells)} setup cells on disk, ready to replay if needed")Store the sandbox id beside the conversation in your own database. That is the only state your servers need to keep.
How do you try two approaches from the same state?
Fork the sandbox. A fork copies the machine as it is, memory included, so each
copy starts with df already loaded and the model's helper functions defined.
Run a different approach in each, compare, keep the winner and stop the rest:
TypeScriptimport { Sandbox } from "withruntime";await using sbx = await Sandbox.create({ memoryMiB: 8192 });await sbx.interpreter.run("import pandas as pd\ndf = pd.read_parquet('/workspace/events.parquet')");const [a, b] = await sbx.fork({ count: 2 }); // each copy has df in memoryconst [dropped, filled] = await Promise.all([ a!.interpreter.run("clean = df.dropna()\nround(clean['value'].mean(), 2)"), b!.interpreter.run( "clean = df.fillna(df.median(numeric_only=True))\nround(clean['value'].mean(), 2)", ),]);console.log(dropped.results[0]?.data["text/plain"], filled.results[0]?.data["text/plain"]);await Promise.all([a!.stop(), b!.stop()]);Pythonfrom withruntime import Sandboxwith Sandbox.create(memory_mib=8192) as sbx: sbx.interpreter.run("import pandas as pd\ndf = pd.read_parquet('/workspace/events.parquet')") a, b = sbx.fork(2) # each copy has df in memory dropped = a.interpreter.run("clean = df.dropna()\nround(clean['value'].mean(), 2)") filled = b.interpreter.run("clean = df.fillna(df.median(numeric_only=True))\nround(clean['value'].mean(), 2)") print(dropped["results"][0]["data"]["text/plain"], filled["results"][0]["data"]["text/plain"]) a.stop() b.stop()A fork's copy is running 2.96 s after the request on Runtime's servers, which beats reloading a large file in each copy. Any outside connection the session held needs reopening in each copy, since two machines cannot share one connection.
The same idea works across conversations. If every conversation starts by loading the same reference data, load it once, take a memory snapshot, and start each new conversation's sandbox from the snapshot: the data is already in memory when the user's first question arrives (snapshots and forks).
Can one sandbox hold several sessions?
Yes. Each interpreter context is its own process with its own variables, in the
same sandbox and on the same disk. That suits one agent with sub-tasks, such
as a planner and a worker that should not overwrite each other's df, or a
notebook interface with several tabs. Create one with
sbx.interpreter.contexts.create({ id: "worker", language: "python" }) and
pass context: "worker" on each run.
Contexts share the sandbox's memory and files, so they are a convenience, not a boundary. Two users, or two customers, get two sandboxes (isolate tenants in a multi-tenant AI app).
In short
- A session's state is variables in memory, files on disk, installed packages and outside connections, and each event touches a different subset.
- Let the model mark setup cells, keep them on the sandbox's disk, and replay
them when
contextStartedorlostsays memory was wiped. - Most deaths are memory: size for three times the largest frame and let the model see its own usage.
- The session lives in the sandbox, so any server can pick up the next message with only the sandbox id.
- Fork to try two approaches from the same loaded state; reopen connections in each copy.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Every Runtime sandbox has a stateful interpreter whose variables survive a pause and are copied by a fork. Read the code interpreter guide, then create an account at withruntime.com and run the recovering session above against your own data.