# Multi-tenant AI apps: how to isolate each customer's code Give every tenant's code its own microVM, tag it with the tenant, and check that tag on every request your backend makes for a user. **Runtime (withruntime.com) runs each sandbox as its own Firecracker microVM with its own kernel, and enforces its CPU, memory, network rules and cost limits on the host, where nothing inside can change them; a new one is running 102 ms after the request on Runtime's servers.** The virtual machine boundary is the part of multi-tenant isolation you can buy. The rest is yours to build: which tenant a sandbox belongs to, what it can reach, whose credentials it holds and how much of your capacity it may use. This post walks through each layer, with the code that enforces it. ## What has to be isolated in a multi-tenant AI app? Seven things, and only the first is solved by choosing a sandbox: | Layer | The failure | Where it is enforced | | ----------- | ----------------------------------------------------------------- | ------------------------------------------------------ | | Compute | Tenant A's code reads tenant B's files or memory | A separate microVM per tenant, at least | | Lookup | User of tenant A passes tenant B's sandbox id to your API | Your backend, on every request | | Network | A's sandbox connects to B's sandbox, or to your internal services | Network rules on the host; sandbox-to-sandbox off | | Credentials | A's code uses a key that belongs to B, or to you | What you put in each sandbox, and what you leave out | | Capacity | A starts so many sandboxes that B's requests wait | Per-tenant quotas in your backend | | Money | A's runaway job eats the budget you planned for everyone | Idle pause and a size per sandbox, a budget per tenant | | Offboarding | A leaves, and their files outlive the contract | One label that finds everything they own | Most incidents in multi-tenant systems happen at the second row, not the first. The machine boundary held; the application handed over the wrong machine. ## One sandbox per tenant, per user or per task? Per user or per task inside the tenant, almost always. A sandbox per tenant shares one machine among all of a company's users, so one user's agent can read another user's files in it. That is acceptable for a team workspace where everyone sees everything anyway, and wrong for most products. | Unit | Isolates | Fits | Watch for | | ---------------- | ----------------------------- | ---------------------------------------------- | ------------------------------------------ | | Per tenant | Companies from each other | Shared team workspaces | Users of one tenant see each other's files | | Per user | Every person from every other | App builders, notebooks, personal agents | Idle machines: let them pause | | Per task | Every job from every other | Code review, CI, data jobs, one-off tool calls | Start time on every task | | Per conversation | Every chat from every other | Chat assistants with a code tool | Name it, so each turn finds it again | The finer units cost little more on Runtime, because a sandbox nobody is using pauses itself after 60 seconds with nothing happening and costs only storage until the next request wakes it. One per user is the default that holds up; [pause between turns](/blog/pause-agent-sandboxes-between-turns) covers the idle side in detail. ## How do you stop a user reaching another tenant's sandbox? Check the tenant on every lookup, in one function that every route goes through. Your backend holds one Runtime key, and that key reaches every sandbox it made, for every tenant. So when a user's request names a sandbox id, the only thing standing between that user and another company's machine is your code. Tag every sandbox with its tenant when you create it, and refuse any id whose tag does not match the caller's tenant: ```ts check import { Runtime, Sandbox } from "withruntime"; const runtime = new Runtime(); const MAX_RUNNING_PER_TENANT = 20; // Create: the tenant and user are labels, set by your backend, never by the user. export async function createFor(tenant: string, user: string) { let running = 0; for await (const _ of await runtime.sandboxes.list({ labels: { tenant }, state: ["running"] })) running++; if (running >= MAX_RUNNING_PER_TENANT) throw new Error(`tenant ${tenant} is at its limit`); return Sandbox.create({ labels: { tenant, user }, idlePauseSeconds: 300, vcpu: 2, // the size comes from the tenant's plan, never from the request memoryMiB: 4096, network: { internet: true, allow: ["pypi.org", "*.pythonhosted.org", "registry.npmjs.org"] }, }); } // Every other route: an id from the user is only a claim until the label agrees. export async function sandboxFor(tenant: string, sandboxId: string) { const sbx = await Sandbox.connect(sandboxId).catch(() => null); if (!sbx || sbx.info.labels.tenant !== tenant) throw new Error("sandbox not found"); return sbx; } const mine = await createFor("acme", "user-17"); console.log("acme reaches", (await sandboxFor("acme", mine.id)).id); console.log( await sandboxFor("globex", mine.id).then( () => "LEAK", (e: Error) => `globex: ${e.message}`, ), ); await mine.stop(); ``` ```python check from withruntime import Runtime, Sandbox runtime = Runtime() MAX_RUNNING_PER_TENANT = 20 def create_for(tenant: str, user: str): """The tenant and user are labels, set by your backend, never by the user.""" running = sum(1 for _ in runtime.sandboxes.list(labels={"tenant": tenant}, state=["running"])) if running >= MAX_RUNNING_PER_TENANT: raise RuntimeError(f"tenant {tenant} is at its limit") return Sandbox.create( labels={"tenant": tenant, "user": user}, idle_pause_seconds=300, vcpu=2, # the size comes from the tenant's plan, never from the request memory_mib=4096, network={"internet": True, "allow": ["pypi.org", "*.pythonhosted.org", "registry.npmjs.org"]}, ) def sandbox_for(tenant: str, sandbox_id: str): """Every other route: an id from the user is only a claim until the label agrees.""" try: sbx = Sandbox.connect(sandbox_id) except Exception: sbx = None if sbx is None or sbx.info["labels"].get("tenant") != tenant: raise LookupError("sandbox not found") return sbx mine = create_for("acme", "user-17") print("acme reaches", sandbox_for("acme", mine.id).id) try: sandbox_for("globex", mine.id) print("LEAK") except LookupError as e: print("globex:", e) mine.stop() ``` Three details matter. The error for a wrong tenant is the same as for a missing id, so a probing user learns nothing. The labels come from your session, never from the request body. And `sandboxFor` is the only way any route gets a sandbox; a code review rule that refuses `Sandbox.connect` anywhere else keeps it that way. Labels are also how you list, count and bill per tenant ([label and list sandboxes](/how-to/label-and-list-sandboxes)). ## What can a tenant's sandbox reach on the network? Only what you allow, and by default not each other. Out of the box a Runtime sandbox cannot connect to private or internal addresses, so your VPC, your cloud's metadata endpoint and other sandboxes are out of reach. Two settings change that, and in a multi-tenant app both deserve a deliberate decision: - **Sandbox-to-sandbox networking** lets every sandbox of your account reach every other by name. It is off until you turn it on, and it is account-wide. For tenant workloads, leave it off; if your own services need it, run them in a separate Runtime account from tenant code ([reach other sandboxes by name](/docs/networking#reach-your-other-sandboxes-by-name)). - **A tunnel into your network** gives your machines a route to sandboxes. Keep tenant sandboxes off any path that leads back into your own systems. Within those, an allow list per sandbox narrows outbound traffic to the package registries and APIs the tenant's code needs, and the rules apply to root inside the sandbox too ([allow only some hosts](/how-to/allow-only-some-hosts)). ## How do you give each tenant its own credentials? Pass them per command, or have the sandbox prove who it is. Do not store a tenant's credential as a Runtime secret: account secrets are available, as a placeholder that works on their hosts, to every sandbox in the account. That is right for something every sandbox may use, and wrong for tenant A's GitHub token, which tenant B's code could then use against GitHub. Two patterns work: 1. **Per command.** Your backend fetches the tenant's token and passes it in that command's `env`. It exists in the sandbox only while the command runs and in whatever the command writes down. Runtime never echoes `env` values back. 2. **Per sandbox, proven.** Code in the sandbox asks Runtime for a signed identity token and sends it to your API. Your API verifies it, reads the sandbox's id from it, looks up that sandbox's tenant label itself, and returns short-lived, tenant-scoped credentials. The tenant comes from Runtime's record, not from anything the sandbox says ([identity tokens](/docs/identity-tokens)). ```js // Your API, with the jose library: which sandbox is calling, according to Runtime? import { createRemoteJWKSet, jwtVerify } from "jose"; const keys = createRemoteJWKSet(new URL("https://withruntime.com/oidc/jwks")); export async function callingSandbox(bearer) { const { payload } = await jwtVerify(bearer, keys, { issuer: "https://withruntime.com/oidc", audience: "https://api.example.com", }); return payload.sandbox_id; // then read its tenant label, mint credentials for that tenant only } ``` The order matters: the token proves which sandbox is calling, and only your own lookup of that sandbox's label says which tenant it serves. A tenant field in the request body would be the caller's claim again. ## What about the apps tenants' agents build? Serve them from Runtime's domain, never yours. A preview of a port in a sandbox gets its own hostname under `runtimehost.com`, so an app a tenant's agent wrote can never read your product's cookies or call your API with a signed-in user's session. If the preview sits in an iframe inside your product, name the tenant's own origin with `embedOrigins`, so no other site can frame it, and keep previews private, with tokens, unless the tenant publishes on purpose ([share a port](/how-to/share-a-port-with-a-preview)). ## How do you keep one tenant from using everyone's capacity? Give each tenant a quota below your account's. Your Runtime account runs 100 sandboxes at once, and every tenant draws from that pool. Without a per-tenant cap, one customer's batch job of a few hundred tasks fills it and everyone else's requests wait. The running count in `createFor` above is that cap. Set it from the tenant's plan, and queue their excess rather than refusing it. Money works the same way. Idle pause and a size you choose bound what one sandbox spends while nobody uses it; summing `chargedMicros` across a tenant's label bounds the tenant, and the [task budget pattern](/blog/agent-cloud-bill-limits) applies to a tenant unchanged. A daily spending limit on your backend's key, which only a person can change, bounds the whole account's day ([set a daily spending limit](/how-to/set-a-daily-spending-limit)). Within a single machine, the CPU, memory and disk you give a sandbox are limits enforced on the host, so a tenant's busy sandbox cannot take more than the size you gave it. For work that must never wait for CPU, [reserve it](/how-to/reserve-cpu). ## How do you remove a tenant's data when they leave? By label, in one pass over everything the tenant created: ```ts import { Runtime } from "withruntime"; const runtime = new Runtime(); const tenant = "acme"; const { stopped, failed } = await runtime.sandboxes.stopAll({ labels: { tenant } }); for await (const sbx of await runtime.sandboxes.list({ labels: { tenant }, includeStopped: true })) await sbx.delete(); console.log( `stopped ${stopped.length}, failed ${failed.length}; deleted every sandbox labelled ${tenant}`, ); ``` ```python from withruntime import Runtime runtime = Runtime() tenant = "acme" result = runtime.sandboxes.stop_all(labels={"tenant": tenant}) for sbx in runtime.sandboxes.list(labels={"tenant": tenant}, include_stopped=True): sbx.delete() print(f"stopped {len(result['stopped'])}, failed {len(result['failed'])}; " f"deleted every sandbox labelled {tenant}") ``` Label volumes and snapshots with the tenant too, and delete them in the same pass. Then a deletion request from a customer is one function and one entry in your own records, not a search. ## In short - The microVM boundary keeps tenants' code apart; your backend keeps their requests apart, by checking a tenant label on every sandbox lookup. - Use a sandbox per user, task or conversation; per tenant only for shared team workspaces. - Leave sandbox-to-sandbox networking off for tenant code, and keep tenants' own credentials out of account-wide secrets. - Cap each tenant's running sandboxes and spend below your account's, and label everything so a tenant can be removed in one pass. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Each tenant's sandbox on Runtime is a Firecracker microVM with network rules and cost limits enforced outside it, running 102 ms after the create request on Runtime's servers. [Open an account at withruntime.com](/sign-in) and put `sandboxFor` in front of your first route.