Throttled requests now return 429, with rate-limit headers
Rate-limited requests now respond with 429 Too Many Requests instead of
503 Service Unavailable, and carry a Retry-After header telling you exactly
how long to wait.
If your client branches on 503 to detect throttling, update it to 429.
What changed
Previously a throttled request came back as a bare 503 with no indication of
when to try again, which is indistinguishable from a genuine outage — so most
HTTP clients treated it as a transient server fault and retried straight back
into the throttle window.
A throttled response now looks like this:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 37
RateLimit-Limit: 300
RateLimit-Remaining: 0
RateLimit-Reset: 37
{
"errors": [
{
"status": 429,
"code": "rate_limited",
"detail": "Rate limit exceeded. Retry after 37 seconds."
}
]
}
Because the limit uses a fixed window and Retry-After gives the exact number
of seconds left in it, you no longer need to guess with exponential backoff —
wait Retry-After seconds and retry.
Documented limits
API-key traffic now has a published ceiling:
| Scope | Limit |
|---|---|
| Per API key | 300 requests / minute |
| Per IP address | 600 requests / minute |
The per-key budget is shared between the REST endpoints and the GraphQL endpoint. The per-IP limit is a backstop set well above the per-key limit; a single well-behaved key never reaches it.
The /mcp endpoint keeps its own separate, lower limits — see
the MCP overview.
What to do
- Branch on
429, not503, to detect throttling. - Honour
Retry-Afterrather than backing off blindly. - Don't expect
RateLimit-*headers on successful responses — they are sent only when you are throttled.
Full details, including copy-pasteable retry helpers in Python and JavaScript, are in Errors & rate limits.