What is an API rate limit?
An API rate limit caps how many requests a client may make in a stretch of time, and answers 429 Too Many Requests past it.
Runtime lets each key make about 50 requests a second, in bursts of up to
200, with 128 served at once, and a request past that waits its turn for up to
five seconds before it is refused. A refusal is a 429 with Retry-After, it
costs nothing, and both SDKs retry it with the same idempotency key
(limits).
The status code
RFC 6585, published in April 2012, defines it: "The 429 status code indicates that the user has sent too many requests in a given amount of time ("rate limiting")." The response "MAY include a Retry-After header indicating how long to wait before making a new request". A client that honours that header, and backs off when it is missing, recovers without anyone noticing.
Kinds of limit
| Limit | Counts | Example on Runtime |
|---|---|---|
| Rate | Requests per second | About 50 a second per key |
| Burst | Requests allowed above the rate at once | Up to 200 |
| Concurrency | Requests in flight at the same moment | 128 per key, 192 per organization |
| Quota | Resources held, not requests | 100 sandboxes at once on a paid account |
A quota is a different refusal: past it a create answers quota_exceeded,
names the limit, and says when a new account's lifts
(how many at once).
Why agents meet rate limits
An agent fans out: forty tool calls in a second is ordinary. Two habits keep it
inside the limits without slowing it down. Retry 429 and 503 after
Retry-After with the same idempotency key, so a
retried create never makes a second sandbox. And let the SDK wait for room when
every slot is taken, rather than failing the task
(handle no capacity).
Related
- What is an idempotency key?
- How to retry safely with idempotency keys
- How to handle no capacity
- What is a webhook?
Sources
Checked 27 September 2026.