See exactly what your AI agent did: commands, files and network
Record every command and its output in your own code, diff the files before and after, and log network destinations outside the sandbox.
Runtime (withruntime.com) keeps its own record of every sandbox: each command's program, exit code and duration for 14 days, every lifecycle event for 14 days, and an account audit log kept at least 400 days that names the sandbox when code inside one calls the API. That record is deliberately thin: it never stores a command's arguments, because arguments carry secrets. When you need to answer "what exactly did the agent do between 2:10 and 2:40?", the full answer comes from three things you record yourself, at the boundary where the agent cannot reach them. This post builds all three.
What can you find out about an agent's run, and from where?
Each question has one best source, and most of them sit outside the sandbox:
| Question | Source | Kept | Can the agent alter it? |
|---|---|---|---|
| Which commands ran, in full? | Your wrapper around exec |
As long as you keep it | No |
| What did each command print? | The same wrapper | As long as you keep it | No |
| Which programs ran, and did they fail? | Runtime's command record | 14 days | No |
| Which files were added, changed, removed? | A manifest before and after, from your code | As long as you keep it | No |
| Which hosts did it connect to? | Your own upstream proxy's log | Your proxy's retention | No |
| Which hosts was it refused? | The allow list, and the error it got | In your command log | No |
| When did it start, pause, stop? | Runtime's events and webhooks | 14 days | No |
| Did it change keys, rules or previews? | Runtime's audit log | At least 400 days | No |
| How busy was the machine? | Runtime's metrics | See the guide | No |
The pattern is the point: nothing in the right-hand column depends on a file inside the sandbox.
Why not keep the log inside the sandbox?
Because the agent is root in there. Anything written inside the sandbox, from
shell history to an audit daemon's file, can be edited or deleted by the
next command, and a model following an injected instruction will happily do
that. Shell history is incomplete anyway: commands run through bash -c or a
script never reach it.
Record where the agent has no hands. Your code calls exec, so your code sees
every command and its output on the way back. Runtime's host enforces the
network rules, so it sees every connection. Neither is inside the machine the
agent controls (root inside the sandbox).
How do you record every command with its output?
Route every command through one function that writes a line to your own log before returning the result. Add a file manifest on either side of the task, and you have the core of a flight recorder:
TypeScriptimport { createHash } from "node:crypto";import { appendFile } from "node:fs/promises";import { Sandbox } from "withruntime";const LOG = "agent-run.jsonl";const hash = (s: string) => createHash("sha256").update(s).digest("hex").slice(0, 16);const clip = (s: string) => (s.length > 4000 ? `${s.slice(0, 2000)}\n[...]\n${s.slice(-2000)}` : s);const record = (entry: object) => appendFile(LOG, JSON.stringify({ at: new Date().toISOString(), ...entry }) + "\n");// Every command the agent runs goes through here, never through sbx.exec directly.export async function recordedExec(sbx: Sandbox, command: string, cwd = "/workspace") { const started = Date.now(); const r = await sbx.exec(command, { cwd, timeoutMs: 300_000 }); await record({ sandbox: sbx.id, kind: "command", command, cwd, exit: r.exitCode, timedOut: r.timedOut, ms: Date.now() - started, stdout: clip(r.stdout), stderr: clip(r.stderr), stdoutHash: hash(r.stdout), }); return r;}// Every file under a folder, as path -> sha256, skipping .git and node_modules.async function manifest(sbx: Sandbox, root = "/workspace") { const r = await sbx.exec( "find . -xdev -type f -not -path './.git/*' -not -path '*/node_modules/*' -print0 | xargs -0 -r sha256sum", { cwd: root, timeoutMs: 120_000 }, ); const lines = r.stdout.split("\n").filter((line) => line.length > 66); return new Map(lines.map((line) => [line.slice(66), line.slice(0, 64)] as const));}export async function recordFileChanges(sbx: Sandbox, before: Map<string, string>) { const after = await manifest(sbx); const added = [...after.keys()].filter((p) => !before.has(p)); const removed = [...before.keys()].filter((p) => !after.has(p)); const changed = [...after] .filter(([p, h]) => before.has(p) && before.get(p) !== h) .map(([p]) => p); await record({ sandbox: sbx.id, kind: "files", added, removed, changed }); return { added, removed, changed };}await using sbx = await Sandbox.create({ labels: { run: "audit-demo" } });const before = await manifest(sbx);await recordedExec(sbx, "echo 'first note' > notes.txt && python3 --version");console.log(await recordFileChanges(sbx, before));Pythonimport hashlibimport jsonimport timefrom datetime import datetime, timezonefrom withruntime import SandboxLOG = "agent-run.jsonl"def record(entry: dict) -> None: with open(LOG, "a") as f: f.write(json.dumps({"at": datetime.now(timezone.utc).isoformat(), **entry}) + "\n")def clip(s: str) -> str: return s if len(s) <= 4000 else s[:2000] + "\n[...]\n" + s[-2000:]def recorded_exec(sbx, command: str, cwd: str = "/workspace"): """Every command the agent runs goes through here, never through sbx.exec directly.""" started = time.monotonic() r = sbx.exec(command, cwd=cwd, timeout_ms=300_000) record({"sandbox": sbx.id, "kind": "command", "command": command, "cwd": cwd, "exit": r.exit_code, "timed_out": r.timed_out, "ms": round((time.monotonic() - started) * 1000), "stdout": clip(r.stdout), "stderr": clip(r.stderr), "stdout_hash": hashlib.sha256(r.stdout.encode()).hexdigest()[:16]}) return rdef manifest(sbx, root: str = "/workspace") -> dict[str, str]: """Every file under a folder, as path -> sha256, skipping .git and node_modules.""" r = sbx.exec("find . -xdev -type f -not -path './.git/*' -not -path '*/node_modules/*' " "-print0 | xargs -0 -r sha256sum", cwd=root, timeout_ms=120_000) return {line[66:]: line[:64] for line in r.stdout.splitlines() if len(line) > 66}def record_file_changes(sbx, before: dict[str, str]) -> dict: after = manifest(sbx) changes = {"added": [p for p in after if p not in before], "removed": [p for p in before if p not in after], "changed": [p for p, h in after.items() if p in before and before[p] != h]} record({"sandbox": sbx.id, "kind": "files", **changes}) return changeswith Sandbox.create(labels={"run": "audit-demo"}) as sbx: before = manifest(sbx) recorded_exec(sbx, "echo 'first note' > notes.txt && python3 --version") print(record_file_changes(sbx, before))A few choices in this code are worth copying even if you write your own:
- Output is clipped, but hashed whole. The head and tail of each stream are what a person reads; the hash proves which full output it was if you keep the raw text elsewhere.
- The environment is never logged. Pass secrets in
envand they stay out of your log as they stay out of Runtime's. Log the names of the variables if you need to know which were set. - One sandbox id per line. When a run forks or uses several sandboxes, the log still reads as one timeline.
Which files did the agent change?
The manifest diff answers it for the folders you list, by content rather
than by timestamp. A timestamp-based find -newer is cheaper but root can
reset a file's time with touch, and a content hash cannot be faked that way.
Two places deserve a manifest beyond the project folder. /etc shows changed
system configuration, such as a new package source or a hosts file entry.
/usr/local shows tools installed outside the package manager. For a
running view while the agent works, a file watch that sends to your webhook
reports writes, renames and deletes under a folder within seconds, without
keeping a sandbox awake (watch files for changes).
The before-and-after manifest is the record; the watch is the live feed.
Which hosts did the agent connect to?
Narrow first, then log. With an allow list on the sandbox, every connection to a host not on the list is refused on the host, and the command that tried it fails with an error your command log captures (allow only some hosts). That covers the "tried to reach somewhere it should not" case without any extra system.
For a record of every destination, paid accounts can send their sandboxes'
connections through their own HTTP proxy after Runtime's rules allow them. Your
proxy can log each connection's host, port, time and byte counts. It cannot read
the encrypted traffic, which is usually what you want: you learn that the
agent talked to api.github.com for 40 KB, not the token it sent
(your own proxy).
TypeScriptimport { Runtime } from "withruntime";const runtime = new Runtime();await runtime.networks.upstreamProxy.set({ url: "http://egress-log.example.com:3128", secret: "PROXY_AUTH",});console.log(await runtime.networks.upstreamProxy.get());If the proxy is down, connections fail rather than going around it, so the log has no gaps an agent could slip through.
What does Runtime record without any setup?
Four things, all readable from code and none writable by the agent:
- Commands: each command's program name, exit code and duration, for 14
days, on the sandbox's page. A failing
pytestat 2:31 shows up there even if your own log was lost. - Events: every create, pause, wake and stop, with the reason a sandbox stopped, for 14 days, and sent to your webhook as they happen.
- The audit log: changes to keys, network rules, secrets, preview visibility and every sandbox created, with who did it, for at least 400 days. A call made by code inside a sandbox names that sandbox, so an agent that was given a key and used it to change network rules is on record with the machine it acted from (read the audit log).
- Metrics: CPU and memory over time, which show when a "quiet" agent was in fact compiling for twenty minutes (read sandbox metrics).
Your flight recorder and Runtime's record should agree. Comparing the two is a cheap integrity check: a command in Runtime's record that is missing from yours means something ran outside your wrapper.
What should a person read after the run?
A one-screen summary built from the log, not the log itself. Nobody reads two hundred JSON lines after every task, but everyone reads five numbers and a list. Generate it at the end of each run and attach it to the pull request or the ticket the agent worked on:
textRun 2026-10-08 14:10-14:41 UTC, sandbox 7f3a2c10, task "fix leap-year parsing"Commands: 48 (3 failed: pytest x2, pip install x1), longest 2 min 14 s (pytest)Files: 4 changed, 1 added (tests/test_leap.py), 0 removed; /etc unchangedNetwork: pypi.org, files.pythonhosted.org; 1 refused (paste.example.com)Account: no key, rule or preview changesThe refused host is the line to read first. An agent that tried to reach a paste site during a bug fix was probably following text it found in an issue or a dependency's README, and that is worth a look whatever the tests say (prompt injection to code execution). Keep the full log beside the summary for as long as you keep the change; a question about a line of code six months later is answered by the command that wrote it.
In short
- Never trust a log the agent could write to; record at your
execcall and at the network edge. - Wrap every
execto log the full command, exit code, duration and clipped output, and keep secrets inenvso they never reach the log. - Diff a content-hash manifest before and after each task; add
/etcand/usr/localwhen the agent hassudo. - An allow list records refusals for free; your own upstream proxy records every destination.
Run it on Runtime
Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Runtime keeps each command's program, exit code and duration for 14 days and its audit log for at least 400 days, outside the reach of the code in the sandbox. Read the observability guide, then create an account at withruntime.com and run the recorder above on your agent's next task.