Rate limits
Credits a minute per tenant, refilling continuously.
| Plan | Credits a minute | Roughly |
|---|---|---|
| Free | 120 | 2 credits a second |
| Build | 600 | 10 credits a second |
| Scale | 3,000 | 50 credits a second |
| Archive | 12,000 | 200 credits a second |
The limit is per tenant, not per key. Issuing a second key does not buy a second allowance — keys exist for revocation and attribution.
Credits, not calls
The minute counts the same credits the month does. A hybrid search takes 3 of it, a story expansion 2, a relationship call up to 16. A call is admitted at its worst case, and whatever it did not use goes back the moment it finishes.
A call larger than a whole minute's credits is not refused forever: it runs once the bucket is full. A 16-credit graph call on a key allowed 12 credits a minute waits until the minute has refilled, then runs.
Free calls
Calls that cost no credits — /account, account_status, every MCP message including the
handshake, and estimate=true previews — take one message token each instead, from a separate allowance of
as many messages a minute as the plan has credits. A request refused before it ran, such as a 400, takes
one too. Past it, the refusal is 429 message_rate_limited.
Reading the headers
x-credits-per-minute: 600
x-credits-per-minute-remaining: 595
x-credits-per-minute-reset: 60 remaining is after this call, with its unused credits returned. reset is
60 on a call that ran, and that is not a bug: the bucket refills continuously rather than at a
wall-clock boundary, so the useful figure is the refill period. On a refusal it is the seconds until the refused call
would fit. A free call carries only x-credits-per-minute.
When you hit it
HTTP/1.1 429 Too Many Requests
retry-after: 2
x-credits-per-minute: 600
x-credits-per-minute-remaining: 0
x-credits-per-minute-reset: 2
x-credits-charged: 0
{
"error": "credit rate limited",
"code": "credit_rate_limited",
"detail": "this minute's credits are spent; the call needs 16 free, which they will be in `retry_after_secs`",
"retry_after_secs": 2,
"credits": { "min": 2, "max": 16 },
"credits_needed": 16,
"credits_per_minute": 600,
"credits_per_minute_remaining": 0
} A correct client
async function call(url: string, key: string): Promise<Response> {
for (let attempt = 0; ; attempt++) {
const response = await fetch(url, { headers: { "x-api-key": key } });
if (response.status !== 429) return response;
// Three conditions share this status, and the body says which.
const { code } = await response.clone().json();
if (code === "insufficient_credits") {
// Not transient: the monthly credits are spent on a key that stops there.
throw new Error("out of credits until the monthly reset; upgrade or wait");
}
if (attempt >= 5) return response;
const wait = Number(response.headers.get("retry-after") ?? 1);
await new Promise((r) => setTimeout(r, wait * 1000));
}
}
Switching on code is the part people leave out. credit_rate_limited and
message_rate_limited are worth waiting for; insufficient_credits is not.
Credits and overage covers it.
Staying under it
- Price before a burst.
estimate=truecosts nothing and tells you what a call will take. - Pick the cheap form. Keyword ranking for a name or an exact phrase,
sort=recencyon/storieswhen you want the newest,evidence=0on a relationship call when you only need the shape. Planning a credit budget lists them. - Ask for bigger pages.
limit=100instead oflimit=10is one tenth of the calls for the same results. - Serialise, do not fan out. Twenty concurrent calls empty a per-minute bucket in the first second. A small concurrency limit with a queue behind it beats a burst plus retries.
- Cache what does not change. An article fetched by id is immutable; there is no reason to fetch it twice.
How it is enforced
The API runs several replicas behind a load balancer, and the bucket is shared across them, so the published figure is the figure — not the figure multiplied by however many replicas happen to be running. If the shared counter is ever unavailable, the limit degrades to being enforced per replica (looser, never stricter) rather than failing requests.