GET /stories
Search, collapsed into deduplicated stories.
One real event across many outlets, with an article count, an outlet count and a representative headline. Use this whenever results go into a context window. Equivalent to the `list_stories` MCP tool. Credits: 1, or 2 semantic and 3 hybrid unless sort=recency; +1 with near, bbox, event_type or amount bounds; +1-2 per entity_id; +1 when signal_percentile_min resolves a threshold; +1 per full 100 rows of offset, which stops at 1000. Pass estimate=true to get the exact range without running the call, or max_credits to refuse anything costlier.
How clustering works
Articles are clustered offline, not at query time, so the same event always carries the same
story_id whichever query surfaced it. Clustering runs in two passes: URL canonicalisation and
near-duplicate detection catch syndication and light rewrites, then a second pass over the entities and topics in
the text catches independent coverage of the same event that shares no wording.
That means the counts are meaningful. count is how many articles are in the cluster;
outlets is how many distinct publisher domains. A story with 40 articles across 3 outlets is one wire
story republished; 40 articles across 35 outlets is widely distributed. Widely distributed is not the same as
independently reported: a wire service or an ownership group can put one newsroom's report on many domains, so
read sources before calling it independent coverage. Those are different facts and
it is worth reading them separately.
Parameters
The same query, filter, signal and time parameters as /search — clustering
changes how results are grouped, not what matches, so a query that worked there works here unchanged, and
mode ranks the clusters the same way it ranks articles (the default is keyword). The
differences are few. Facets count over articles, so they belong on the article endpoints; a story row is already
the compact shape, so there is no view; cursor pages a newest-first list of articles, so
it is not here either. sort is the one parameter this endpoint takes that /search does
not.
- Ranking
-
qmode - Filters
-
sourcesource_typepublisher_countrylanguagestory_idtopicorganizationpersoncountryauthorlocationlocation_idcityregionnamementioned_datequote_verbamount_objectestimatemax_credits - Entity ids
-
entity_id - Place
-
nearradius_kmbbox - Event classes
-
event_type - Amounts
-
amount_minamount_max - Signals
-
signalsignal_minsignal_maxsignalssignal_percentile_min - Time range
-
fromto - Paging, shape and facets
-
offsetlimit
Each group is described once, on the page linked beside it, rather than restated here — the descriptions are generated from the same OpenAPI document either way.
The geographic, event-class and
amount filters narrow stories as they narrow articles: a story is listed when one
of its articles in the window matches. A story row has no matched_locations. To see which cities
matched, pass the row's story_id and the same area to /search.
Any of those filters adds 1 credit to credits_charged: the query that checks whether the index holds the
field.
Sort
| Name | Type | Description |
|---|---|---|
| sort | `count` · `outlets` · `recency` | Story order: `count` (articles, default), `outlets` (distinct publishers), or `recency` (newest representative article first, ranked newest-first before collapsing). |
sort decides which events come first. count, the default, puts the stories with the most
articles first: the loudest events, syndication included. outlets puts the most widely carried first,
so one wire story republished forty times on three domains drops below an event thirty outlets covered once each.
recency puts the newest first. It ranks articles newest-first and then collapses them, so each row's
representative article is the story's newest member in the window. That is the order a feed wants: what is new,
one row per event.
Response
| Field | Type | Description |
|---|---|---|
| attribution | string | Required source attribution for any use of these results. |
| coverage | `complete` · `timed_out` · `range_too_wide` · `recent_unavailable` · `field_unavailable` | Why the answer is or is not complete, named for what you can do about it. `complete`: everything in range was read (or provably could not have changed the page). `timed_out`: the query ran out of budget -- retry, or narrow the range. `range_too_wide`: the range spans more of the corpus than one request may read -- narrow it; retrying unchanged will not help. `recent_unavailable`: the most recent data in range could not be read -- retry shortly; widening the range will not help. `field_unavailable`: a filter needs a field part of the index was built without (geographic, event-class and amount fields arrive with each shard's rebuild) -- the answer covers only the rebuilt part; retrying before the rebuild will not help. |
| credits_charged | integer | Index queries this answer took, and what the call is charged (also the `x-credits-charged` header). One for a plain search; more with `entity_id`, `signal_percentile_min`, a cursor page, a geographic or amount filter (a query per batch of 500 candidates checked), or a derived-field filter on stories, semantic or hybrid search and aggregate (one query to check the index holds the field). |
| from_clamped | boolean | Whether a `from` you supplied was moved forward to that boundary. Distinguishes a plan limit from a corpus with no coverage. |
| has_more | boolean | Whether another page exists. Stop paging when false. |
| history_days | integer | How far back the caller's plan may query; null = the full archive. |
| limit | integer | The page size that was applied. |
| offset | integer | The zero-based offset this page starts at. |
| partial | boolean | Whether this answer is incomplete. `true` means part of the index could not be read, so an empty or short result set is NOT evidence of absence -- retry rather than caching it as a negative. |
| resolved | ResolvedEntity[] | Present when `entity_id` was used: the spellings each id matched. |
| signal_threshold | SignalThreshold | Present when `signal_percentile_min` was used: the threshold it resolved to. |
| stories | Story[] | This page of stories, largest first. |
| total | integer | Matches in the whole set (see `total_relation`), not on this page. |
| total_relation | `exact` · `approximate` | `exact` for keyword and metadata queries, and for geographic and amount filters when every candidate was checked; `approximate` for ANN-derived modes, which rank a bounded candidate set, and for a checked filter whose candidates exceeded the scan. `/stories` is always `approximate`: stories are collapsed from a bounded ranked set of articles. |
Story
| Field | Type | Description |
|---|---|---|
| count | integer | Articles in the cluster. |
| outlets | integer | Distinct publisher domains in the cluster. Domains, not newsrooms: a wire story reaches many domains from one report. |
| published_at | string | `published_at` of the representative article: the newest member under `sort=recency`, the best-ranked otherwise. |
| sources | string[] | Publisher domains in the cluster, capped; `outlets` is the full count. |
| story_id | string | Stable cluster id; pass to /stories/{id} or back as the `story_id` filter. |
| title | string | Headline of the representative article. |
| url | string | URL of the representative article. |
title and url are a representative article from the cluster, not a synthesised summary —
the API never generates text. published_at is when that article was published: the newest member
under sort=recency, the best-ranked otherwise. It is not when the story began;
first_seen on /stories/{id} is. A row is deliberately thin: when it was first and last seen, every outlet, and what
it is about are one request away.
Expanding a story
Pass a story_id to /stories/{id} for the whole event
in one request: the counts, every outlet, first_seen and last_seen, the organizations,
people and topics across its members, and its newest articles as compact hits.
# One event, then everything about it.
curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=grid outage" --data-urlencode "limit=5"
curl -s "https://api.unzoi.com/stories/s-9f2c?limit=10" -H "x-api-key: $UNZOI_KEY"
When the members themselves are the point — paging through all of them, faceting them by publisher country,
filtering them to one language — pass the id back to /search as the
story_id filter instead:
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "story_id=s-9f2c" \
--data-urlencode "facets=publisher_country,language" \
--data-urlencode "limit=50" That two-step is the usual shape for an agent: cluster to decide what happened, expand to see how it was covered. The guide works through it.
When not to use this
When coverage is the question. "How many outlets ran this?", "did the framing differ by country?", "find me
every mention of this company" — those want articles, and collapsing them throws away the answer. Use
/search with facets instead.