Runtime

Five ways AI agents run up your cloud bill, and the limit for each

An agent runs up a cloud bill through forgotten machines, runaway loops, busy processes, piled-up storage and traffic out.

Runtime (withruntime.com) bills an agent sandbox on the CPU it actually uses, $0.025 per active vCPU-hour, and pauses it after 60 seconds with nothing happening, so the most common runaway, a machine nobody is using, costs almost nothing by default. The other four need a limit you set. This post walks through each way an agent spends money it should not, works out the worst case for a day, and shows the limit that stops it, including which limits an agent can quietly raise and the one it never can.

Why are agents worse at spending than people?

Because an agent never gets the feeling that something is taking too long. A person notices a laptop fan, a slow terminal or a bill alert. An agent retries what failed, starts another sandbox when one is busy, and keeps going until something stops it. Every failure below is ordinary agent behavior, not a bug, and each has been the cause of a surprise bill somewhere.

The worst cases in this post use a 2 vCPU, 4 GiB sandbox, which costs $0.08 an hour with both CPUs working and $0.03125 an hour while it waits, at Runtime's published rates (pricing).

1. What does a forgotten sandbox cost?

On most clouds, a full day of running time; on Runtime, a minute of it, then storage. An agent creates a sandbox, the task ends or the process crashes, and nothing ever calls stop(). The machine keeps its memory reserved and the meter running.

  • Without a limit: 24 hours waiting cost $0.75, and a fleet of a hundred forgotten ones $75 a day.
  • With the default idle pause: after 60 seconds with no request, no running command, no open connection and no CPU use, the sandbox pauses itself and pays only for its saved state. Two GB of it cost $0.0053 a day.

The limit: keep idlePauseSeconds low, from 10 seconds up, and create sandboxes in a scope that deletes them: await using in TypeScript, with in Python. A label on every sandbox a task creates lets a sweeper find what the scope missed (pause when idle).

2. What happens when an agent loops on create?

It fills your account. An agent whose tool call times out tries again; a planner that thinks a sandbox is busy starts another; a framework retries a failed step with a fresh machine each time. Each attempt is a new sandbox, and each one is busy installing the same dependencies.

  • Without a limit: a paid account runs 100 sandboxes at once. All of them busy for a day cost $192.
  • With limits: an idempotency key made from the task id makes a retried create answer the sandbox it already made instead of a second one. And the size of each sandbox is a constant in your code, not a tool argument, so an agent cannot ask for 16 vCPUs and 64 GiB because a build was slow. A daily spending limit on the agent's key (below) caps the loop you did not foresee.

The Runtime SDKs send an idempotency key with every write already, so a retry inside one process never duplicates a create; pass your own to cover a process that restarts (retry safely).

3. What does a runaway process cost?

As much as the machine can burn. A test that spins forever, a crawler that never finds its stop condition, a compile in a loop: the sandbox is not idle, so it never pauses, and on a busy machine CPU billing on use stops helping.

  • Without a limit: the largest sandbox, 16 vCPUs and 64 GiB, fully busy for a day costs $21.
  • With limits: timeoutMs on each command ends that command, and timeoutSeconds at create gives the whole sandbox a time limit: when it runs out, the sandbox stops running, its files kept, and compute billing stops with it. A 2 vCPU, 4 GiB sandbox with a 30-minute limit can cost at most $0.04, whatever runs inside it.

Give every command a timeout that fits the step. A test run that normally takes two minutes gets ten, not the 24 hours a streamed command may run for. Give every task sandbox a time limit that fits the task.

4. How does storage pile up?

Quietly, one snapshot at a time. Storage is the bill nobody watches, because no single item is expensive. Say an agent snapshots after each of 50 steps in a task, runs 20 tasks a day, keeps every snapshot 7 days with retentionDays: 7 (with none, a snapshot is kept as long as you have credit), and each stores 2 GB of its own:

TextSnapshots kept:   50 × 20 × 7 days       = 7,000Stored:           7,000 × 2 GB           = 14,000 GBMonthly cost:     14,000 × $0.08 per GB  = $1,120With retentionDays: 1 instead of 7      = $160Deleted when each task ends             ≈ nothing

The limit: pass retentionDays: 1 to checkpoint snapshots, delete a task's snapshots when it finishes, and give volumes a size you chose rather than one the agent guessed, since a volume is charged on its full size from the moment it exists. A block two of your snapshots share counts once, so real chains usually cost less than this sum (storage prices).

5. Can an agent run up a bill with network traffic?

Yes, by sending data out. Inbound traffic is free on Runtime, and each account's first 100 GiB out a month is free; past that it costs $0.02 per GB. An agent that uploads build artifacts in a loop, or a compromised one that exfiltrates a dataset, spends it.

  • Without a limit: a paid sandbox moves at most 500 GiB a day, in and out together. All of it out, past the allowance, costs $10.74 for that one sandbox.
  • With limits: an allow list in the sandbox's network rules narrows it to the hosts the task needs, such as your package registry and your model provider. Data cannot leave for anywhere else, so it cannot be billed either (allow only some hosts).

Which limits can the agent raise by itself?

Every one it can reach with its own key. That is the part most budget designs miss. Network rules can be widened with network.set(), idle pause can be turned off with update(), and a paused sandbox can be woken with a new time limit, all by any key that may run the sandbox. A limit the agent can change is advice.

Limit Stops Set with Who can change it
Idle pause Forgotten sandboxes idlePauseSeconds at create Any key that runs the sandbox
Size in your code Oversized machines vcpu, memoryMiB at create Whoever creates
Command timeoutMs Hung and looping commands Each exec Whoever runs the command
Sandbox time limit One runaway sandbox timeoutSeconds at create Any key that runs the sandbox
Network allow list Traffic out, exfiltration Create or network.set() Any key that runs the sandbox
Snapshot retention Storage pile-up retentionDays Whoever snapshots
Account concurrency Fan-out Your account Support, on request
Daily spending limit on a key Everything that key spends The website, by a person Only a person; no key, ever

Two rules follow. First, keep the Runtime key in your orchestrator, the code that calls the model and runs its tools, and not in the sandbox where the model's commands run. The model then cannot call update() at all. Second, give each agent its own key with a daily spending limit. It counts everything that key's sandboxes cost in any 24 hours, and past it a create or wake fails with spending_limit_reached and charges nothing (set a daily spending limit).

What does the code look like with every limit in place?

Create each task's sandboxes through one function that applies the limits, and check the task's total before every new one. This runs as written:

TypeScriptimport { randomUUID } from "node:crypto";import { Runtime } from "withruntime";const runtime = new Runtime();const TASK_BUDGET_MICROS = 2_000_000; // $2 for everything one task startsasync function spentOn(task: string) {  let micros = 0;  const page = await runtime.sandboxes.list({ labels: { task }, includeStopped: true });  for await (const sbx of page) micros += Number(sbx.info.chargedMicros);  return micros;}async function sandboxFor(task: string, attempt: number) {  const spent = await spentOn(task);  if (spent >= TASK_BUDGET_MICROS) {    await runtime.sandboxes.stopAll({ labels: { task } });    throw new Error(`task ${task} has spent its budget`);  }  return runtime.sandboxes.create(    {      labels: { task },      vcpu: 2, // oversized: the size is yours, never the agent's      memoryMiB: 4096,      idlePauseSeconds: 30, // forgotten: pauses half a minute after the last use      timeoutSeconds: 1800, // runaway: the sandbox stops after 30 minutes      network: { internet: true, allow: ["pypi.org", "*.pythonhosted.org"] }, // traffic out    },    { idempotencyKey: `${task}-${attempt}` }, // a retried create answers the same sandbox  );}const taskId = `task-${randomUUID()}`; // your task's own id, the same on every retryawait using sbx = await sandboxFor(taskId, 1);const run = await sbx.exec("python3 -c 'print(sum(range(10)))'", { timeoutMs: 120_000 });console.log(run.exitCode, run.stdout.trim(), await spentOn(taskId));
Pythonfrom uuid import uuid4from withruntime import Runtimeruntime = Runtime()TASK_BUDGET_MICROS = 2_000_000  # $2 for everything one task startsdef spent_on(task: str) -> int:    page = runtime.sandboxes.list(labels={"task": task}, include_stopped=True)    return sum(int(sbx.info["chargedMicros"]) for sbx in page)def sandbox_for(task: str, attempt: int):    spent = spent_on(task)    if spent >= TASK_BUDGET_MICROS:        runtime.sandboxes.stop_all(labels={"task": task})        raise RuntimeError(f"task {task} has spent its budget")    return runtime.sandboxes.create(        labels={"task": task},        vcpu=2,  # oversized: the size is yours, never the agent's        memory_mib=4096,        idle_pause_seconds=30,  # forgotten: pauses half a minute after the last use        timeout_seconds=1800,  # runaway: the sandbox stops after 30 minutes        network={"internet": True, "allow": ["pypi.org", "*.pythonhosted.org"]},  # traffic out        idempotency_key=f"{task}-{attempt}",  # a retried create answers the same sandbox    )task_id = f"task-{uuid4()}"  # your task's own id, the same on every retrywith sandbox_for(task_id, 1) as sbx:    run = sbx.exec("python3 -c 'print(sum(range(10)))'", timeout_ms=120_000)    print(run.exit_code, run.stdout.strip(), spent_on(task_id))

chargedMicros is what each sandbox has cost so far, in millionths of a dollar, so the task's total includes the sandboxes it already stopped. The check runs before every new sandbox, and each sandbox's time limit bounds what it can add after the check: at most $0.04 for 30 minutes with both CPUs busy. So the task can pass its budget by at most one sandbox's worth, however many it starts.

An agent can also read its key's own limit before it plans a large fan-out, and stop to ask a person instead of being refused halfway:

TypeScriptimport { Runtime } from "withruntime";const runtime = new Runtime();const { daily } = await runtime.limits.get();const left = daily.remainingMicros === null ? Infinity : Number(daily.remainingMicros);if (left < 5_000_000) console.log("Under $5 left today: ask before starting the fan-out.");

How do you catch spending you did not plan for?

Watch lifecycle events, not invoices. A webhook on sandbox.stopped tells you the moment a sandbox reaches its time limit, with stopReason saying why it stopped; a burst of sandbox.created events from one key is a loop in progress. Usage & billing shows what each product cost each day and how long your credit lasts at this month's pace, and every owner is told by email when the balance will last under five days. Credit is prepaid, so the balance never goes below zero and you never owe more than you put in (metrics and webhooks).

In short

  • Forgotten sandboxes are handled by idle pause; keep it on and keep it short.
  • Loops, oversized machines and runaway processes need idempotency keys, a size set in your code, a time limit on every sandbox and a timeout on every command.
  • Storage and traffic out grow quietly; set snapshot retention to a day and narrow the network to the hosts a task needs.
  • An agent holding a key can raise most limits. Keep the key in your orchestrator, and give each agent's key a daily limit only a person can change.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Runtime charges $0.025 per vCPU-hour of CPU actually used and $0.0075 per GiB-hour of memory, with no plan fee, and every limit in this post is built in. Start free, with 100 hours of a 2 vCPU, 4 GB sandbox included every month and no card at withruntime.com, and set the daily spending limit on your first key before an agent ever uses it.

100 hours of a 2 vCPU, 4 GB sandbox,included every month.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card