Runtime

What is an API rate limit?

An API rate limit caps how many requests a client may make in a stretch of time, and answers 429 Too Many Requests past it.

Runtime lets each key make about 50 requests a second, in bursts of up to 200, with 128 served at once, and a request past that waits its turn for up to five seconds before it is refused. A refusal is a 429 with Retry-After, it costs nothing, and both SDKs retry it with the same idempotency key (limits).

The status code

RFC 6585, published in April 2012, defines it: "The 429 status code indicates that the user has sent too many requests in a given amount of time ("rate limiting")." The response "MAY include a Retry-After header indicating how long to wait before making a new request". A client that honours that header, and backs off when it is missing, recovers without anyone noticing.

Kinds of limit

Limit Counts Example on Runtime
Rate Requests per second About 50 a second per key
Burst Requests allowed above the rate at once Up to 200
Concurrency Requests in flight at the same moment 128 per key, 192 per organization
Quota Resources held, not requests 100 sandboxes at once on a paid account

A quota is a different refusal: past it a create answers quota_exceeded, names the limit, and says when a new account's lifts (how many at once).

Why agents meet rate limits

An agent fans out: forty tool calls in a second is ordinary. Two habits keep it inside the limits without slowing it down. Retry 429 and 503 after Retry-After with the same idempotency key, so a retried create never makes a second sandbox. And let the SDK wait for room when every slot is taken, rather than failing the task (handle no capacity).

Sources

Checked 27 September 2026.