Rate limits and backoff
429 means three different things here, and treating them the same way produces either a client that
hangs for a day or one that hammers a wall.
The three 429s
credit_rate_limited | message_rate_limited | insufficient_credits | |
|---|---|---|---|
| Means | This minute's credits are spent | Too many free calls this minute | Not enough credits left this month, on a key that stops there |
retry-after | When this call's credits refill, usually 1–2 seconds | About a second | Seconds until the month resets — possibly days |
| Charged | Nothing | Nothing | Nothing |
| Retrying helps | Yes, soon | Yes, soon | Not the same call; a cheaper one might fit |
The body's code is the discriminator. Every error is the same JSON body, so a correct
client reads it rather than guessing from headers.
Do not back off exponentially from the rate limit
The per-minute bucket refills continuously, not at a wall-clock boundary, and retry-after is
computed from the credits the refused call needs. A client that doubles its way from 1 to 60 seconds spends most of
it idle on allowance it already paid for — and it does this precisely when it is busiest.
Exponential backoff is for 503, where the server is actually unwell and pressure makes it worse.
A correct client
async function call(url, key) {
for (let attempt = 0; ; attempt++) {
const response = await fetch(url, { headers: { "x-api-key": key } });
if (response.status === 429) {
const body = await response.clone().json();
// Not transient. Retrying burns nothing, but it will not work for days.
if (body.code === "insufficient_credits") {
throw new OutOfCredits(body.retry_after_secs, body.credits_remaining);
}
if (attempt >= 5) throw new Error("rate limited repeatedly");
// Transient, and retry-after is the honest wait.
await sleep(Number(response.headers.get("retry-after") ?? 1) * 1000);
continue;
}
// The server, not you. Back off properly here.
if (response.status === 503 && attempt < 3) {
await sleep((2 ** attempt) * 1000 + Math.random() * 500);
continue;
}
return response;
}
}
The jitter on the 503 path matters if you have several workers: without it they retry in lockstep and
arrive together every time.
Not hitting it in the first place
Backoff is the fallback. These are cheaper:
- Bound your concurrency. Twenty parallel calls exhaust a per-minute bucket in the first second and spend the rest of the minute retrying. A semaphore of 4–8 with a queue behind it finishes sooner.
- Spend fewer credits per call. Keyword ranking where a name or phrase is enough,
/storieswithsort=recencyfor the newest,evidence=0on relationship calls when you only need the shape. Planning a credit budget has the list. - Ask for bigger pages.
limit=100is one tenth of the calls oflimit=10for the same results, and each page is charged. - Cache by id. An article fetched with
/doc/{id}is immutable. There is no reason to fetch it twice. - Read the headers you already have. Every response carries
x-credits-per-minute-remaining; a client that slows down at 10% remaining never sees a429.
Pacing from the headers
let pauseUntil = 0;
function pace(response) {
const remaining = Number(response.headers.get("x-credits-per-minute-remaining") ?? Infinity);
const limit = Number(response.headers.get("x-credits-per-minute") ?? 0);
// Below 10% of the bucket, spread the rest of the minute out rather than
// sprinting into a 429 and retrying.
if (limit && remaining < limit * 0.1) {
pauseUntil = Date.now() + 1000;
}
} On MCP
Too many messages is still a 429 with retry-after on the HTTP request — send the same
JSON-RPC message again. A tool call refused for credits is not: it arrives as a tool error carrying
credit_rate_limited or insufficient_credits on an open session, so an agent can wait, or
report it, rather than concluding the server is down. The shapes are here.