Pagination and facets
Paging
Two ways. offset and limit page any query. limit is 1–100 and defaults to
10; offset goes up to 100,000. A newest-first browse can also page by
cursor, which does not drift as new articles arrive.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=inflation" \
--data-urlencode "offset=20" --data-urlencode "limit=20"
Stop on has_more: false rather than by comparing offset + limit to total.
The two agree for keyword queries and can disagree for ranked ones, for the reason below.
view sits beside the paging parameters because it changes how much of each hit comes back rather
than which hits do: compact keeps the id, title, url, source, time, language, story and score. A
hundred compact hits cost the same request as a hundred full ones and far fewer tokens.
The field list.
Deep paging
The offset ceiling is not arbitrary: a large offset makes the server rank and discard everything before it. To walk a large ranked result set, narrow by time instead and page within each window — a month at a time is both faster and cheaper than an offset in the tens of thousands. To walk everything newest-first, use a cursor.
Cursor paging (recency)
A newest-first browse — /search with no q, or
/top-headlines — can be paged by position rather than by offset.
Pass cursor=start for the first page. The response carries next_cursor; pass it back as
cursor for the next page. When next_cursor is absent, there is no further page.
Offsets drift on a feed that is still growing. An article indexed between two requests pushes everything down one
place, and the next page repeats a hit. A cursor names the last article you were handed, so pages are disjoint
however much arrives in between. The order is fixed: published_at newest first, then
id. The tiebreak matters more than it sounds. Articles arrive in batches that share one timestamp, and
a cursor walks a batch of thousands without skipping or repeating one.
# Every article from Singaporean publishers in August, newest first, 100 at a time.
cursor=start
while [ -n "$cursor" ]; do
page=$(curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "publisher_country=SG" \
--data-urlencode "from=2026-08-01" --data-urlencode "to=2026-08-31" \
--data-urlencode "view=compact" --data-urlencode "limit=100" \
--data-urlencode "cursor=$cursor")
# An incomplete page is not a page: stop, and resume from the same cursor.
echo "$page" | jq -e '(.partial | not) and .coverage == "complete"' > /dev/null \
|| { echo "incomplete page; resume with cursor=$cursor" >&2; break; }
echo "$page" | jq -c '.results[]'
cursor=$(echo "$page" | jq -r '.next_cursor // empty')
done - A browse only.
cursorwith aqis refused, whatever themode: a ranked result set has no fixed position to resume from, so it pages byoffset.cursorwithoffsetis refused too. Both are400withcode: invalid_parameter. - No geographic or amount filters.
near,bbox,amount_minandamount_maxare checked candidate by candidate, which a cursor cannot page: withcursorthey are refused with400 invalid_parameter. Page those withoffset. - The same request every page. The cursor is a position, not a saved query. Send the same filters and window with each page.
- Opaque. Pass
next_cursorback as it came. Its format is not part of the contract. - No depth ceiling. Each page is a fresh query below the cursor's position, so the 100,000 offset limit does not apply, and page 500 costs what page 2 does.
A cursor page is charged a credit per index query it takes, reported as credits_charged: up to two for
the first page and up to three for each page after, in place of the page's base price. They are the page itself, the rest of the timestamp the cursor sits in, and the whole
tie group at the page's new boundary, read in full and ordered by id. Across indexes a page is charged
what the busiest one needed, not the sum. That re-read is what makes the order
deterministic. A single timestamp holding more than 5,000 matching articles cannot be ordered within that bound;
the page comes back partial with coverage: range_too_wide, and the fix is a narrower
filter, not a retry.
Totals
total_relation tells you how much to trust total.
| Value | Means | When |
|---|---|---|
exact | total is the true number of matching articles. | Keyword and metadata-only queries, which the inverted index can count exhaustively. A geographic or amount filter is exact when every candidate was checked. |
approximate | total is how many candidates were considered, or an estimate, not how many exist. |
Semantic, hybrid and similarity queries, which rank a bounded candidate set rather
than scoring the whole corpus. Also a keyword or browse query whose geographic or amount filter had more
candidates than the 2,000 it checks: total is estimated from the share that passed.
|
So do not render "About 4,132 results" for a hybrid query — it is a count of what the ANN pass looked at. If you
need a true count, run the same filters as a keyword query and read that total instead.
Facets
Facets count over the whole match set, not over the page you were handed. That is what makes them worth a request: they tell you what exists before you filter on it.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=lithium supply" \
--data-urlencode "facets=source,country,topic" \
--data-urlencode "facet_limit=10" \
--data-urlencode "limit=5" "facets": {
"country": {
"values": [
{ "value": "AU", "count": 214 },
{ "value": "CL", "count": 168 }
],
"sum_other_doc_count": 91,
"doc_count_error_upper_bound": 0,
"approximate": false
}
} Up to 8 facet fields per request, each returning up to 100 values. Any filter field can be faceted.
Reading the error bounds
-
sum_other_doc_count— how many matches fell outside the returned values. A large number meansfacet_limitis cutting off a long tail. -
doc_count_error_upper_bound— the most any returned count could be understated.0means the counts are exact. -
approximate—truewhen that upper bound is non-zero. Counts are exact for keyword queries and can be approximate for ranked ones, for the same reason totals are. A request with a geographic or amount filter marks every facet approximate, because its facets count candidates before each one is checked.
The practical use is to discover rather than to report: facet by source and
country, see who is actually covering something, then re-query filtered to one of them. Presenting a
facet count to a user as a precise figure is only safe when approximate is false.
Evaluating coverage works through it.
Stories page too
/stories takes the same offset and limit and
reports the same total, total_relation and has_more — where
total counts clusters, not articles. It does not take facets, because facet counts are over
articles and belong on the article endpoints, and it does not take cursor: for newest events first,
use sort=recency and page by offset.
/stories/{id} takes limit only. It lists the
cluster's newest members up to that many and says whether more exist (articles_truncated); there is
no offset, because the endpoint is for the story as one object. To walk every member, or to facet
them, use /search?story_id= with the paging above.