Skip to content
unzoi docs
Search and navigation
Start here
REST API
MCP
Limits and plans
Agent clients
SDKs
Guides

Rate limits

Credits a minute per tenant, refilling continuously.

PlanCredits a minuteRoughly
Free 120 2 credits a second
Build 600 10 credits a second
Scale 3,000 50 credits a second
Archive 12,000 200 credits a second

The limit is per tenant, not per key. Issuing a second key does not buy a second allowance — keys exist for revocation and attribution.

Credits, not calls

The minute counts the same credits the month does. A hybrid search takes 3 of it, a story expansion 2, a relationship call up to 16. A call is admitted at its worst case, and whatever it did not use goes back the moment it finishes.

A call larger than a whole minute's credits is not refused forever: it runs once the bucket is full. A 16-credit graph call on a key allowed 12 credits a minute waits until the minute has refilled, then runs.

Free calls

Calls that cost no credits — /account, account_status, every MCP message including the handshake, and estimate=true previews — take one message token each instead, from a separate allowance of as many messages a minute as the plan has credits. A request refused before it ran, such as a 400, takes one too. Past it, the refusal is 429 message_rate_limited.

Reading the headers

x-credits-per-minute: 600
x-credits-per-minute-remaining: 595
x-credits-per-minute-reset: 60

remaining is after this call, with its unused credits returned. reset is 60 on a call that ran, and that is not a bug: the bucket refills continuously rather than at a wall-clock boundary, so the useful figure is the refill period. On a refusal it is the seconds until the refused call would fit. A free call carries only x-credits-per-minute.

When you hit it

HTTP/1.1 429 Too Many Requests
retry-after: 2
x-credits-per-minute: 600
x-credits-per-minute-remaining: 0
x-credits-per-minute-reset: 2
x-credits-charged: 0

{
  "error": "credit rate limited",
  "code": "credit_rate_limited",
  "detail": "this minute's credits are spent; the call needs 16 free, which they will be in `retry_after_secs`",
  "retry_after_secs": 2,
  "credits": { "min": 2, "max": 16 },
  "credits_needed": 16,
  "credits_per_minute": 600,
  "credits_per_minute_remaining": 0
}

A correct client

async function call(url: string, key: string): Promise<Response> {
  for (let attempt = 0; ; attempt++) {
    const response = await fetch(url, { headers: { "x-api-key": key } });
    if (response.status !== 429) return response;

    // Three conditions share this status, and the body says which.
    const { code } = await response.clone().json();
    if (code === "insufficient_credits") {
      // Not transient: the monthly credits are spent on a key that stops there.
      throw new Error("out of credits until the monthly reset; upgrade or wait");
    }

    if (attempt >= 5) return response;
    const wait = Number(response.headers.get("retry-after") ?? 1);
    await new Promise((r) => setTimeout(r, wait * 1000));
  }
}

Switching on code is the part people leave out. credit_rate_limited and message_rate_limited are worth waiting for; insufficient_credits is not. Credits and overage covers it.

Staying under it

  • Price before a burst. estimate=true costs nothing and tells you what a call will take.
  • Pick the cheap form. Keyword ranking for a name or an exact phrase, sort=recency on /stories when you want the newest, evidence=0 on a relationship call when you only need the shape. Planning a credit budget lists them.
  • Ask for bigger pages. limit=100 instead of limit=10 is one tenth of the calls for the same results.
  • Serialise, do not fan out. Twenty concurrent calls empty a per-minute bucket in the first second. A small concurrency limit with a queue behind it beats a burst plus retries.
  • Cache what does not change. An article fetched by id is immutable; there is no reason to fetch it twice.

How it is enforced

The API runs several replicas behind a load balancer, and the bucket is shared across them, so the published figure is the figure — not the figure multiplied by however many replicas happen to be running. If the shared counter is ever unavailable, the limit degrades to being enforced per replica (looser, never stricter) rather than failing requests.