What is CPU overcommitment?
CPU overcommitment is promising more vCPUs to the machines on a host than it has cores, on the bet that they will not all be busy at once.
Runtime sells that bet back to you as a lower price: a shared sandbox is
guaranteed a floor and billed for the CPU it measures, $0.025 per active
vCPU-hour. When a job needs its cores held at all times, cpu: "reserved"
keeps every vCPU for it and bills them as always busy
(pricing).
How it works
A hypervisor schedules vCPUs onto physical cores the way an operating system schedules threads. Most machines are idle most of the time, so a host with 16 cores can run guests with far more than 16 vCPUs between them and still give each one a core when it asks. The ratio of promised vCPUs to cores is the overcommit ratio.
Linux does this with control groups. cpu.weight, from 1 to 10,000 with a
default of 100, sets a group's share when cores are contested, and cpu.max
sets a hard ceiling: "the group may consume up to $MAX in each $PERIOD
duration". Kubernetes exposes the same pair as a container's CPU request and
limit, and says that at a limit "the kernel will restrict access to the CPU".
Overcommitment done openly and quietly
| Approach | What you are told | What you pay for |
|---|---|---|
| Quietly overcommitted VM | "2 vCPUs" | 2 vCPUs, all the time |
| Burstable instance | A baseline and credits | The instance, plus overages |
| Shared CPU on Runtime | A floor, and bursts up to vcpu |
The CPU the code measured |
| Reserved CPU on Runtime | Every vCPU, held | Every vCPU, busy or not |
The honest version names the guarantee. On Runtime that is the
CPU floor: a twentieth of a vCPU by default, raised per
sandbox with cpuFloorMillis.
Why it suits agents
An agent's sandbox waits on a model most of the minute, then compiles or runs tests for a few seconds. Sharing cores across many such sandboxes is what lets the waiting cost $0.03125 an hour at 2 vCPUs and 4 GiB instead of the full $0.08.
Related
Sources
Checked 27 September 2026.