Skip to content
unzoi docs
Search and navigation
Start here
REST API
MCP
Limits and plans
Agent clients
SDKs
Guides

Pagination and facets

Paging

Two ways. offset and limit page any query. limit is 1–100 and defaults to 10; offset goes up to 100,000. A newest-first browse can also page by cursor, which does not drift as new articles arrive.

curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
  --data-urlencode "q=inflation" \
  --data-urlencode "offset=20" --data-urlencode "limit=20"

Stop on has_more: false rather than by comparing offset + limit to total. The two agree for keyword queries and can disagree for ranked ones, for the reason below.

view sits beside the paging parameters because it changes how much of each hit comes back rather than which hits do: compact keeps the id, title, url, source, time, language, story and score. A hundred compact hits cost the same request as a hundred full ones and far fewer tokens. The field list.

Deep paging

The offset ceiling is not arbitrary: a large offset makes the server rank and discard everything before it. To walk a large ranked result set, narrow by time instead and page within each window — a month at a time is both faster and cheaper than an offset in the tens of thousands. To walk everything newest-first, use a cursor.

Cursor paging (recency)

A newest-first browse — /search with no q, or /top-headlines — can be paged by position rather than by offset. Pass cursor=start for the first page. The response carries next_cursor; pass it back as cursor for the next page. When next_cursor is absent, there is no further page.

Offsets drift on a feed that is still growing. An article indexed between two requests pushes everything down one place, and the next page repeats a hit. A cursor names the last article you were handed, so pages are disjoint however much arrives in between. The order is fixed: published_at newest first, then id. The tiebreak matters more than it sounds. Articles arrive in batches that share one timestamp, and a cursor walks a batch of thousands without skipping or repeating one.

# Every article from Singaporean publishers in August, newest first, 100 at a time.
cursor=start
while [ -n "$cursor" ]; do
  page=$(curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
    --data-urlencode "publisher_country=SG" \
    --data-urlencode "from=2026-08-01" --data-urlencode "to=2026-08-31" \
    --data-urlencode "view=compact" --data-urlencode "limit=100" \
    --data-urlencode "cursor=$cursor")
  # An incomplete page is not a page: stop, and resume from the same cursor.
  echo "$page" | jq -e '(.partial | not) and .coverage == "complete"' > /dev/null \
    || { echo "incomplete page; resume with cursor=$cursor" >&2; break; }
  echo "$page" | jq -c '.results[]'
  cursor=$(echo "$page" | jq -r '.next_cursor // empty')
done
  • A browse only. cursor with a q is refused, whatever the mode: a ranked result set has no fixed position to resume from, so it pages by offset. cursor with offset is refused too. Both are 400 with code: invalid_parameter.
  • No geographic or amount filters. near, bbox, amount_min and amount_max are checked candidate by candidate, which a cursor cannot page: with cursor they are refused with 400 invalid_parameter. Page those with offset.
  • The same request every page. The cursor is a position, not a saved query. Send the same filters and window with each page.
  • Opaque. Pass next_cursor back as it came. Its format is not part of the contract.
  • No depth ceiling. Each page is a fresh query below the cursor's position, so the 100,000 offset limit does not apply, and page 500 costs what page 2 does.

A cursor page is charged a credit per index query it takes, reported as credits_charged: up to two for the first page and up to three for each page after, in place of the page's base price. They are the page itself, the rest of the timestamp the cursor sits in, and the whole tie group at the page's new boundary, read in full and ordered by id. Across indexes a page is charged what the busiest one needed, not the sum. That re-read is what makes the order deterministic. A single timestamp holding more than 5,000 matching articles cannot be ordered within that bound; the page comes back partial with coverage: range_too_wide, and the fix is a narrower filter, not a retry.

Totals

total_relation tells you how much to trust total.

ValueMeansWhen
exact total is the true number of matching articles. Keyword and metadata-only queries, which the inverted index can count exhaustively. A geographic or amount filter is exact when every candidate was checked.
approximate total is how many candidates were considered, or an estimate, not how many exist. Semantic, hybrid and similarity queries, which rank a bounded candidate set rather than scoring the whole corpus. Also a keyword or browse query whose geographic or amount filter had more candidates than the 2,000 it checks: total is estimated from the share that passed.

So do not render "About 4,132 results" for a hybrid query — it is a count of what the ANN pass looked at. If you need a true count, run the same filters as a keyword query and read that total instead.

Facets

Facets count over the whole match set, not over the page you were handed. That is what makes them worth a request: they tell you what exists before you filter on it.

curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
  --data-urlencode "q=lithium supply" \
  --data-urlencode "facets=source,country,topic" \
  --data-urlencode "facet_limit=10" \
  --data-urlencode "limit=5"
"facets": {
  "country": {
    "values": [
      { "value": "AU", "count": 214 },
      { "value": "CL", "count": 168 }
    ],
    "sum_other_doc_count": 91,
    "doc_count_error_upper_bound": 0,
    "approximate": false
  }
}

Up to 8 facet fields per request, each returning up to 100 values. Any filter field can be faceted.

Reading the error bounds

  • sum_other_doc_count — how many matches fell outside the returned values. A large number means facet_limit is cutting off a long tail.
  • doc_count_error_upper_bound — the most any returned count could be understated. 0 means the counts are exact.
  • approximate — true when that upper bound is non-zero. Counts are exact for keyword queries and can be approximate for ranked ones, for the same reason totals are. A request with a geographic or amount filter marks every facet approximate, because its facets count candidates before each one is checked.

The practical use is to discover rather than to report: facet by source and country, see who is actually covering something, then re-query filtered to one of them. Presenting a facet count to a user as a precise figure is only safe when approximate is false. Evaluating coverage works through it.

Stories page too

/stories takes the same offset and limit and reports the same total, total_relation and has_more — where total counts clusters, not articles. It does not take facets, because facet counts are over articles and belong on the article endpoints, and it does not take cursor: for newest events first, use sort=recency and page by offset.

/stories/{id} takes limit only. It lists the cluster's newest members up to that many and says whether more exist (articles_truncated); there is no offset, because the endpoint is for the story as one object. To walk every member, or to facet them, use /search?story_id= with the paging above.