Runtime

12 questions to ask any agent sandbox provider, with Runtime's answers

Ask a sandbox provider about isolation, measured speed, idle cost, pause, network, keys, limits, forks, images, uptime and switching.

Runtime (withruntime.com) answers each question below with a figure you can check: a Firecracker microVM per sandbox, a first command 221 ms after the create request on Runtime's servers, and $0.025 per vCPU-hour of CPU actually used. Every provider's landing page says fast, secure and cheap. These twelve questions get past the adjectives. Each one comes with the follow-up that separates a real answer from a rehearsed one, and a way to check it yourself in a few minutes.

1. Does every sandbox get its own kernel?

It should, because code that shares a kernel with other tenants is one kernel bug away from them.

Follow up: "What exactly is shared between two customers' sandboxes on the same server?" A container shares the host kernel. A user-space kernel shares less. A microVM shares only the hypervisor. Check it: get the isolation technology named in writing, not just "isolated". Inside a sandbox, uname -r shows the kernel your code talks to. Runtime: every sandbox is a Firecracker microVM with its own Linux kernel (microVM vs container).

2. How fast is a new sandbox, and how was that measured?

Ask for a median and a 95th percentile, timed from the request to the first command finishing, because "ready" can mean many things.

Follow up: "Is that timed on your servers or from a client? Is it the API answering, or a command having run?" A create call that returns before the machine can run anything moves the wait into your first command. Check it: time a create and a first command a hundred times from where your agent runs, and look at the slowest five. Runtime: 221 ms at the median and 484 ms at p95 on Runtime's servers, and 331 ms from a laptop in the US Mountain time zone, network included, with the method published (speed).

3. What does an hour of waiting cost?

Less than an hour of working, if the provider bills CPU by use. Agents spend most of their time waiting on a model, so this question decides the bill.

Follow up: "Do you bill CPU on what the sandbox uses, or on what it was allocated, for every second it runs?" Then ask about minimums per sandbox and plan fees. Check it: price one of your real runs at the provider's published rates: running time, CPU-seconds used, memory held. Runtime: a 2 vCPU, 4 GiB sandbox costs $0.03125 an hour while it waits and $0.08 with both CPUs busy, with no plan fee. On the pricing guide's example job it costs 42% to 88% less than fourteen other providers (what 1,000 agent runs cost).

4. Does pause keep memory, and how fast is the wake?

A pause that keeps only files makes the next turn rebuild every running process; one that keeps memory makes the gap invisible.

Follow up: "Does a paused sandbox wake by itself on the next command, and how long is paused state kept?" Check it: start a process that holds a value in memory, pause, wake, and ask the process for the value. A clean boot that looks fine is not proof. Runtime: pause keeps files, memory and processes; a paused sandbox is running again 76 ms after the wake request on Runtime's servers, and any command or file call wakes it on its own. Paid accounts keep paused state while they have credit (pause, don't delete).

5. What can a sandbox reach on the network?

By default, nothing private, and you should be able to narrow the rest to a list.

Follow up: "Where is the rule enforced? Can root inside the sandbox change it?" A rule enforced inside the machine is a rule the agent can remove. Check it: from inside, request the cloud metadata address, a private address and a host that should be off your list. Runtime: a proxy on the host carries every connection; private and internal addresses are refused for every sandbox, and an allow list narrows the rest, which root inside cannot change (why the sandbox shouldn't reach the whole internet).

6. Can code use an API key without being able to read it?

It should, because an agent that can read a key can be talked into printing it.

Follow up: "Is the secret injected as an environment variable, or added to requests outside the sandbox?" The first is convenient storage; only the second keeps the value out of the agent's reach. Check it: store a secret, then run env in the sandbox and look for it. Runtime: a secret appears as a placeholder, and the host's proxy swaps in the value only on HTTPS requests to the hosts you named (secrets sandboxes never see).

7. How long can one sandbox run?

As long as the work takes. A hard cap on lifetime forces you to checkpoint and restart long jobs at the provider's convenience, not yours.

Follow up: "What happens at the limit: a stop with warning, or a kill? Is the limit on running time or on wall-clock time, pauses included?" Check it: read the create call's options for a maximum, then run something for longer than an hour. Runtime: no time limit unless you set one (timeoutSeconds); a sandbox runs while it works and pauses itself after 60 seconds with nothing happening.

8. How large can one sandbox be, and how many run at once?

Large enough for your heaviest task, and enough of them for your busiest hour.

Follow up: "Is concurrency a hard cap, or do extra creates wait? Can a sandbox have a GPU?" Check it: start as many as your peak needs at once and count how many are running a minute later. Runtime: up to 16 vCPUs and 64 GiB of memory per sandbox, and 100 sandboxes at once on a paid account; creates beyond that wait for room rather than fail. Runtime runs on CPUs: there are no GPU sandboxes, so GPU work belongs elsewhere.

9. Can you copy a running sandbox?

You should be able to, memory included, because best-of-N attempts, evals and tree search all start from one prepared machine.

Follow up: "Does a fork copy running processes, or only the disk? How many copies per call?" Check it: start a server in a sandbox, fork it, and send a request to the copy without starting anything. Runtime: one fork call returns up to 100 running copies with files, memory and processes; a single copy is running 2.96 s after the request on Runtime's servers (best-of-N coding agents).

10. Can you bring your own image?

You should be able to start from any Dockerfile or registry image, so the agent does not spend its first minute installing your stack.

Follow up: "Are images versioned, and can a sandbox switch images without losing its files?" Check it: build your real project's image and time a create from it. Runtime: custom images from a package list, any public or private registry image, or a Dockerfile, versioned and tagged, with start and ready commands (images).

11. What is promised when the service is down?

A number in writing, measured from outside, with something you get back when it is missed.

Follow up: "Is the credit automatic, or do I have to file a claim? Where is uptime published?" Check it: read the terms, then find the status page and its history. Runtime: paid accounts are promised 99% API uptime each month, checked from outside every two minutes; a month below it returns 10% of that month's charges as credit, without asking (uptime promise, status).

12. How hard is it to leave, or to arrive?

It should be a day's work either way. Plain Linux, plain HTTP and standard images keep it that way.

Follow up: "Is there anything in my code that only your platform can run?" Check it: list every provider call in your code; count the ones with no equivalent elsewhere. Runtime: sandboxes are ordinary Linux machines and images are Dockerfiles. Code written for E2B, Daytona, Vercel Sandbox or Blaxel runs after changing one import (switch to Runtime).

How do you check five of these in a minute?

Run one script against each candidate. This one times the first command from your side of the network, then asks the sandbox about its kernel, its size and what it can reach:

TypeScriptimport { Sandbox } from "withruntime";const started = performance.now();await using sbx = await Sandbox.create();await sbx.exec("true");console.log(`first command: ${Math.round(performance.now() - started)} ms, network included`);const checks: Record<string, string> = {  kernel: "uname -r",  "vCPUs and memory": "nproc && free -m | awk '/Mem:/ {print $2 \" MiB\"}'",  "who am I": "id -un && sudo -n true && echo 'sudo works'",  "cloud metadata": "curl -sf -m 5 http://169.254.169.254/ || echo refused",  "private network": "curl -sf -m 5 http://10.0.0.1/ || echo refused",};for (const [name, command] of Object.entries(checks)) {  const result = await sbx.exec(command, { timeoutMs: 15_000 });  console.log(`${name}: ${result.stdout.trim().replaceAll("\n", ", ")}`);}
Pythonimport timefrom withruntime import Sandboxstarted = time.perf_counter()with Sandbox.create() as sbx:    sbx.exec("true")    print(f"first command: {(time.perf_counter() - started) * 1000:.0f} ms, network included")    checks = {        "kernel": "uname -r",        "vCPUs and memory": "nproc && free -m | awk '/Mem:/ {print $2 \" MiB\"}'",        "who am I": "id -un && sudo -n true && echo 'sudo works'",        "cloud metadata": "curl -sf -m 5 http://169.254.169.254/ || echo refused",        "private network": "curl -sf -m 5 http://10.0.0.1/ || echo refused",    }    for name, command in checks.items():        result = sbx.exec(command, timeout_ms=15_000)        print(f"{name}: {', '.join(result.stdout.strip().splitlines())}")

Run it ten times and keep the slowest first command, not the fastest. Port the five checks to each provider's SDK; the shell commands stay the same.

What does the scorecard look like?

One line per question, with the answer to hold out for:

# Question The answer to hold out for Runtime
1 Own kernel per sandbox? Yes, a microVM Firecracker microVM
2 Speed, and how measured? Median and p95 to a first command 221 ms median, on servers
3 Cost of an idle hour? Far below a busy hour $0.03125 idle, $0.08 busy
4 Pause keeps memory? Yes, with a wake on demand Yes, wakes in 76 ms
5 Network rules where? Outside the sandbox On the host; root cannot change them
6 Keys the code cannot read? Injected outside the sandbox Placeholder, value added by the proxy
7 Longest run? No cap you did not set None unless you set one
8 Size and concurrency? Your peak, with room 16 vCPUs, 100 at once
9 Copy a running sandbox? Memory and processes included Up to 100 copies per call
10 Your own image? Any Dockerfile or registry image Yes, versioned
11 Uptime in writing? A number and automatic credit 99%, 10% back
12 Easy to leave? Plain Linux, standard images One import to switch from E2B and others

The guide to choosing an agent sandbox turns these into a full evaluation plan, with the failure cases worth testing.

In short

  • Ask for medians and p95s, timed to a first command, and check them from where your agent runs.
  • Ask whether CPU is billed on use; an agent mostly waits, and waiting should be cheap.
  • Ask where network rules and secrets are enforced; inside the sandbox is not enough.
  • Ask what pause keeps and what a fork copies; memory is the difference.
  • Get uptime and switching terms in writing before you depend on either.

Run it on Runtime

A Runtime sandbox is a Firecracker microVM that runs its first command 221 ms after the create request on Runtime's servers, costs $0.03125 an hour while it waits at 2 vCPU and 4 GiB, and pauses itself with its memory kept. Start with 100 free hours, no card: sign in and run the script above, or read get started and the pricing.

Your first 100 hoursare on us.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Claim 100 hours free