Runtime

What happens when your AI agent runs rm -rf: four setups compared

It deletes whatever the command can reach: on a laptop, your work; in a sandbox with no credentials, a copy you rebuild in seconds.

Runtime (withruntime.com) gives each agent its own Firecracker microVM with its own kernel, so a deleting command reaches that machine and what you attached to it, and a fresh copy from a snapshot is running 440 ms after the request on Runtime's servers. Every team that runs coding agents eventually has the conversation: what if it runs rm -rf? The honest answer is that it will, sooner or later, and not out of malice. The useful question is what the command can reach when it does. This post walks through four common setups, what one destructive command does in each, and why a list of forbidden commands protects less than it seems.

Why would an agent run rm -rf at all?

Almost always as a cleanup that goes wrong, the same way it goes wrong for people:

  • An empty variable. rm -rf "$BUILD_DIR/"* with BUILD_DIR unset becomes rm -rf /*. Steam's Linux client shipped exactly this bug in January 2015 and deleted users' home directories.
  • The wrong directory. The model runs cd build in one command and rm -rf * in the next, but each command starts in the original working directory.
  • Starting fresh. A model stuck on a broken dependency tree decides to delete node_modules, the lockfile, and then the cache folder it thinks is in the project but is in your home directory.
  • Instructions it read. A README, an issue or a web page that says "run this cleanup script first" (prompt injection to code execution).

None of these needs a bad model. They need a shell, a path and one wrong assumption, and an agent issues hundreds of commands a day.

What does it delete in each setup?

Same command, four places to run it:

Setup rm -rf ~ deletes What else the agent can reach Recovery
On your laptop Your home: every repo, your SSH keys Your cloud logins, your browser's sessions From your backups, if you have them; hours
A container with the project mounted in The container's home, and the project through the mount Whatever else is mounted, often the Docker socket Committed work from git; the rest is gone
A cloud sandbox holding your credentials The sandbox's files only Every system its tokens open The sandbox is rebuilt; the remote damage is not
A sandbox with no credentials, short allow list, snapshot A disposable copy Little: the hosts on its list A new sandbox from the snapshot, in 440 ms

The first row is why people worry. An agent running on a developer's laptop has the developer's reach: every repository, every key in ~/.ssh, every cloud CLI that is logged in.

The second row is the one that surprises. A container feels isolated, but a bind mount is your real directory, not a copy. rm -rf /workspace inside the container deletes the folder on your disk, uncommitted work included. Mount the Docker socket for convenience and the container can start a new container with your whole disk mounted.

The third row moves the risk rather than removing it. The machine is disposable, so rm -rf itself is harmless. But the same class of mistake with a credential in reach is not: git push --force with a write token, aws s3 rm --recursive with real keys, DROP TABLE with a production connection string. The delete happens on a system that has no undo.

The fourth row is the goal: a machine of its own, credentials it cannot read, a network that reaches only what the task needs, and a saved state to return to.

Can't you just block rm -rf?

You can block the string. You cannot block the intent, because a shell offers dozens of ways to delete a directory, and a model looking for a way around a refused command will find one without meaning any harm. Here is a typical guard, and six commands that do the same thing as the one it blocks:

TypeScriptimport { Sandbox } from "withruntime";const GUARD = /\brm\s+-[a-z]*(rf|fr)/i; // the usual patternconst sameEffect = [  "rm -rf /workspace/project",  "rm -r -f /workspace/project",  "rm --recursive --force /workspace/project",  "find /workspace/project -delete",  "python3 -c \"import shutil; shutil.rmtree('/workspace/project')\"",  "perl -MFile::Path -e 'rmtree(\"/workspace/project\")'",  "echo cm0gLXJmIC93b3Jrc3BhY2UvcHJvamVjdA== | base64 -d | sh",];// In a sandbox, letting each one run costs nothing: there is nothing here to lose.await using sbx = await Sandbox.create({ labels: { demo: "rm-rf" } });for (const command of sameEffect) {  await sbx.exec("mkdir -p /workspace/project && echo keep > /workspace/project/notes.txt");  const blocked = GUARD.test(command);  if (!blocked) await sbx.exec(command, { cwd: "/workspace", timeoutMs: 30_000 });  console.log(blocked ? "blocked" : "ran    ", command);}
Pythonimport refrom withruntime import SandboxGUARD = re.compile(r"\brm\s+-[a-z]*(rf|fr)", re.I)  # the usual patternsame_effect = [    "rm -rf /workspace/project",    "rm -r -f /workspace/project",    "rm --recursive --force /workspace/project",    "find /workspace/project -delete",    "python3 -c \"import shutil; shutil.rmtree('/workspace/project')\"",    "perl -MFile::Path -e 'rmtree(\"/workspace/project\")'",    "echo cm0gLXJmIC93b3Jrc3BhY2UvcHJvamVjdA== | base64 -d | sh",]# In a sandbox, letting each one run costs nothing: there is nothing here to lose.with Sandbox.create(labels={"demo": "rm-rf"}) as sbx:    for command in same_effect:        sbx.exec("mkdir -p /workspace/project && echo keep > /workspace/project/notes.txt")        blocked = bool(GUARD.search(command))        if not blocked:            sbx.exec(command, cwd="/workspace", timeout_ms=30_000)        print("blocked" if blocked else "ran    ", command)

One line is blocked and six run. A longer pattern catches more of them and still misses a script the model writes to a file and then runs. Guards like this are worth having as a hint to the model, which reads the refusal and usually picks a safer path, but they are not a boundary. The boundary has to be what the machine can reach, not what the command says.

What does a sandbox change, and what doesn't it?

It changes the reach, not the behavior. The agent still runs the command, and the command still deletes everything it can. What a sandbox controls is the list of things it can:

  • Its own disk. A microVM has its own kernel and filesystem. Nothing on your laptop or on another customer's machine is underneath it.
  • What you attach. A volume or a mounted bucket is reachable like any folder, so attach only what the task needs, and mount a bucket with readOnly when the agent should only read it. A volume is backed up off its server every day, which is a recovery point, not a reason to let an agent write to one it does not need (storage and backups).
  • Credentials. On Runtime, a secret reaches the sandbox as a placeholder, and the real value is added on the way out, only for the hosts you named. The agent can use the token for its task and cannot read it, print it or send it elsewhere (secrets).
  • The network. An allow list limits which systems a command can touch at all, so a deleting command aimed at a host you did not list fails to connect (allow only some hosts).

Root inside the sandbox is real root over that machine, and that is fine. The agent can apt-get install what it needs, and the worst it can do with that power is break a machine you were going to throw away (root inside the sandbox).

How do you get back to where you were?

Take the saved state before the risky part, not after the damage. A snapshot of the sandbox after setup, with dependencies installed and the repository cloned, turns any disaster into one call: start a new sandbox from it, which is running 440 ms after the request, and continue. For finer steps, git inside the sandbox gives you undo per command, and a fork lets the agent try a dangerous fix on a copy first. Both are in an undo button for AI agents.

The time limit is part of recovery too. A sandbox created with timeoutSeconds pauses on its own when the limit is up, even if the agent's loop has wedged after breaking something, and its files are kept so you can look at what happened.

What should still need a person?

Anything that leaves the sandbox and cannot be taken back. Inside the machine, let the agent work; at the edge, put a check:

Action Who approves Why
Edit, build, test, delete files Nobody The machine is disposable
Install packages The allow list Only the registries you listed are reachable
Open a pull request Your code, after tests pass Reviewable and reversible
Push to main, deploy, migrate A person Hard or impossible to undo
Write to a production database A person, or never No snapshot covers someone else's system

The pattern from a test gate for agent-written code fits here: the agent's work leaves the sandbox as a diff, and your code, not the agent, decides whether it goes anywhere.

In short

  • An agent will eventually run a destructive command by accident; plan for what it can reach, not for whether it happens.
  • On a laptop it reaches your work and logins; in a container, everything mounted into it; with credentials, every system they open.
  • Command guards are a hint to the model, not a boundary: one line in seven was blocked above.
  • Give each agent its own microVM, placeholder secrets, a short allow list and a snapshot to return to.
  • Keep irreversible actions outside the sandbox behind your code or a person.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). A Runtime sandbox is a Firecracker microVM with its own kernel, its secrets as placeholders the agent cannot read, and a snapshot to return to in 440 ms. Read how security works, then create an account at withruntime.com and run the six commands above where they cannot hurt anything.

400 sandbox hours,every month.

Hours of a 1 GB sandbox, included free.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card