GET /search
Search articles.
Ranked articles with topics, entities, locations and signals. Equivalent to the `search_news` MCP tool. Credits: 1 for keyword ranking or a browse, 2 for semantic, 3 for hybrid; 1-4 with near, bbox, amount_min or amount_max, whose candidates are checked in batches of 500; a cursor page costs 1-2 to start and 1-3 after; +1-2 per entity_id; +1 when signal_percentile_min resolves a threshold; +1 per full 1000 rows of offset. Pass estimate=true to get the exact range without running the call, or max_credits to refuse anything costlier.
Ranking
| Name | Type | Description |
|---|---|---|
| q | string | Query text. Omit for a recency browse. |
| mode | `keyword` · `semantic` · `hybrid` · `recent` | Ranking mode. `hybrid` is usually right when a model wrote the query, because models paraphrase. `recent` orders newest-first and is implied when `q` is omitted. |
mode is the parameter worth thinking about. Use hybrid when a model wrote the query —
models paraphrase, and keyword ranking cannot bridge "central bank tightening" to an article that says "Fed raises
rates". Use keyword when a human gave you an exact phrase, a product name or a proper noun they want
matched literally. The guide goes deeper.
Omitting q entirely turns this into a recency browse — mode is reported as
recent, and you can ask for it by name — which is what
/top-headlines does explicitly. The same is true of the
search_news tool: q is optional there too, so
search_news { story_id: … } lists a story's members with no query at all.
Filters
Exact matches on indexed attributes. Adding words to q makes ranking fuzzier; adding a filter does not,
and costs nothing.
| Name | Type | Description |
|---|---|---|
| source | string | Source domain, e.g. bbc.co.uk. |
| source_type | `web` · `citation` · `academic_archive` · `defense_archive` · `journal_archive` · `non_textual` · `other` | The kind of source: web | citation | academic_archive | defense_archive | journal_archive | non_textual | other. |
| publisher_country | string | The publisher's canonical two-character country identifier. |
| language | string | Source language, ISO-639-3, e.g. eng, fra, ara, zho. |
| story_id | string | A cluster id from /stories, to expand one story into its articles. |
| topic | string | An exact canonical topic identifier returned by the API. |
| organization | string | An exact organization entity. |
| person | string | An exact person entity. |
| country | string | A mentioned country's canonical two-character identifier. |
| author | string | An exact article author. |
| location | string | An exact mentioned location. |
| location_id | string | An exact canonical location identifier. |
| city | string | An exact mentioned city. |
| region | string | An exact mentioned region. |
| name | string | An exact proper name. |
| mentioned_date | string | An exact date mentioned in the article text. |
| quote_verb | string | An exact verb introducing a quotation. |
| amount_object | string | An exact object described by a numeric amount. |
| estimate | boolean | Return this call's price instead of running it: the credit range, a breakdown, your remaining credits and whether the call would run. A preview costs 0 credits. |
| max_credits | integer | Refuse the call, at no cost, when its worst case would cost more than this many credits. The refusal carries the estimate. |
Values are canonical identifiers, not free text — take them from a previous response, resolve a name with
/entities, or read them off a facet
count. Filters and signals explains the vocabularies. A parameter name that
is not in these tables is refused with 400 and code: unknown_parameter, not ignored.
Entity ids
| Name | Type | Description |
|---|---|---|
| entity_id | string | Up to five entity ids from resolve_entity, comma separated (over REST the parameter may also repeat). Each matches any spelling in its alias group in the requested window; several ids must all match. The spellings used are reported in `resolved`. |
An id from /entities matches every spelling in its alias group, where a
filter matches one. The response's resolved lists the spellings each id used, and resolving them is
counted in credits_charged. Filters and signals has the rules.
Place, event classes and amounts
| Name | Type | Description |
|---|---|---|
| near | string | A point, `latitude,longitude` in decimal degrees, e.g. 48.8566,2.3522. With radius_km: articles mentioning a city within that distance. On article results (/search, /top-headlines) each hit reports the cities that matched as matched_locations. |
| radius_km | number | With near: the distance in kilometres, 0.1-2000. |
| bbox | string | A box, `south,west,north,east` in decimal degrees: articles mentioning a city inside it. West greater than east crosses the antimeridian. Not with near. |
| event_type | `bankruptcy` · `earnings_report` · `ipo` · `cyber_attack` · `outage` · `industrial_accident` · `sanctions` · `trade_dispute` · `boycott` · `nationalization` · `privatization` · `money_laundering` · `executive_change` · `factory_closure` · `supply_shortage` · `labor_strike` · `antitrust_action` | An event class derived from the article's GKG themes (events-v1). Precise rather than complete: an article can describe an event without carrying its class. |
| amount_min | number | Articles mentioning a numeric amount of at least this value. With amount_max, one amount must satisfy both; with amount_object, that amount must describe the object. |
| amount_max | number | Articles mentioning a numeric amount of at most this value. With amount_min, one amount must satisfy both; with amount_object, that amount must describe the object. |
These filter on fields the index derives from each article's own records: the cities it mentions, the event classes
its themes imply, and the amounts it mentions. Geographic search covers near,
radius_km and bbox; event classes and
amount bounds are in Filters and signals. A geographic or amount filter is
checked against each candidate article, which can make total_relation approximate and adds to
credits_charged: Exactness and cost.
Signals
Every article carries versioned industry, business-context and risk-context intensities. Range-filtering on them finds coverage with a particular character rather than a particular word.
| Name | Type | Description |
|---|---|---|
| signal | `economy` · `healthcare` · `agriculture` · `labor` · `environment` · `energy` · `transportation` · `real_estate` · `finance` · `defense` · `science_technology` · `trade` · `public_sector` · `financial_uncertainty` · `financial_negative` · `financial_positive` · `legal_litigation` · `financial_stability_stress` · `anxiety` · `conflict` · `supply_disruption` · `cyber_incident` · `climate` · `health_security` · `governance_risk` | A normalized business signal to range-filter on, e.g. finance, energy, conflict. An unknown name is rejected. |
| signal_min | number | Inclusive lower bound for `signal`. |
| signal_max | number | Inclusive upper bound for `signal`. |
| signals | string | Several signal ranges at once, ANDed: `name[:min[:max]]`, comma separated, e.g. `finance:1.5:,energy::2`. |
| signal_percentile_min | number | With `signal`: keep only articles at or above this percentile of the signal's intensity over the match set (50-99.9). The threshold used is reported as `signal_threshold`. Not with `signal_min`. |
Use signal + signal_min/signal_max for one constraint,
signals for several at once, ANDed, or signal + signal_percentile_min for
the top slice of whatever matched, with the value it resolved to reported as signal_threshold. See
Entity and signal filtering.
Time range
| Name | Type | Description |
|---|---|---|
| from | string | Lower time bound, e.g. 2026-07-09 or 20260709120000. Clamped forward to what the caller's plan may reach; see `history_days` and `from_clamped` in the response. |
| to | string | Upper time bound, same formats as `from`. |
Paging, shape and facets
| Name | Type | Description |
|---|---|---|
| offset | integer | Zero-based result offset, up to 100000. |
| limit | integer | Results per page, 1-100 (default 10). |
| cursor | string | Keyset paging for a newest-first browse (no `q`): pass `start` for the first page, then the `next_cursor` the response carried. Pages are deterministic and disjoint even across articles sharing one timestamp. The first page costs two index queries and each later page three. Not with `offset`, `q`, or a ranked `mode`. |
| facets | string | Comma-separated fields to count over the full match set, up to 8. |
| facet_limit | integer | Values per facet, 1-100 (default 10). |
| view | `full` · `compact` | How much of each hit to return. `full` (default) is the whole Article; `compact` keeps id, title, url, source, published_at, language, story_id and score. Use compact when results feed a model. |
offset pages any query. cursor pages a browse with no q, newest first: pass
start, then each next_cursor, and the pages stay disjoint while new articles arrive. The
two do not combine. Pagination and facets covers cursor paging, what the counts mean,
and when they are approximate.
view changes the shape of a hit, not the match set. compact keeps
id, language, published_at, score, source, story_id, title, url and drops the rest: the entity lists, the
signal map, the tone scores. Use it when results feed a model that is deciding what to read next — the full record
is one /doc/{id} away — and leave it at full when your
code reads the fields.
Response
| Field | Type | Description |
|---|---|---|
| attribution | string | Required source attribution for any use of these results. |
| coverage | `complete` · `timed_out` · `range_too_wide` · `recent_unavailable` · `field_unavailable` | Why the answer is or is not complete, named for what you can do about it. `complete`: everything in range was read (or provably could not have changed the page). `timed_out`: the query ran out of budget -- retry, or narrow the range. `range_too_wide`: the range spans more of the corpus than one request may read -- narrow it; retrying unchanged will not help. `recent_unavailable`: the most recent data in range could not be read -- retry shortly; widening the range will not help. `field_unavailable`: a filter needs a field part of the index was built without (geographic, event-class and amount fields arrive with each shard's rebuild) -- the answer covers only the rebuilt part; retrying before the rebuild will not help. |
| credits_charged | integer | Index queries this answer took, and what the call is charged (also the `x-credits-charged` header). One for a plain search; more with `entity_id`, `signal_percentile_min`, a cursor page, a geographic or amount filter (a query per batch of 500 candidates checked), or a derived-field filter on stories, semantic or hybrid search and aggregate (one query to check the index holds the field). |
| facets | object | Present only when `facets` was requested; keyed by the public field name. |
| from_clamped | boolean | Whether a `from` you supplied was moved forward to that boundary. Distinguishes a plan limit from a corpus with no coverage. |
| has_more | boolean | Whether another page exists. Stop paging when false. |
| history_days | integer | How far back the caller's plan may query; null = the full archive. |
| limit | integer | The page size that was applied. |
| mode | `keyword` · `semantic` · `hybrid` · `recent` · `similar` | The ranking that was applied. `recent` when no `q` was given; `similar` from /similar/{id}. |
| next_cursor | string | On a cursor-paged browse: pass back as `cursor` for the next page. Absent when there is no further page or `cursor` was not used. |
| offset | integer | The zero-based offset this page starts at. |
| partial | boolean | Whether this answer is incomplete. `true` means part of the index could not be read, so an empty or short result set is NOT evidence of absence -- retry rather than caching it as a negative. |
| resolved | ResolvedEntity[] | Present when `entity_id` was used: the spellings each id matched. |
| results | Article[] | This page of hits, best first. |
| signal_threshold | SignalThreshold | Present when `signal_percentile_min` was used: the threshold it resolved to. |
| total | integer | Matches in the whole set (see `total_relation`), not on this page. |
| total_relation | `exact` · `approximate` | `exact` for keyword and metadata queries, and for geographic and amount filters when every candidate was checked; `approximate` for ANN-derived modes, which rank a bounded candidate set, and for a checked filter whose candidates exceeded the scan. `/stories` is always `approximate`: stories are collapsed from a bounded ranked set of articles. |
Two pairs of fields describe the answer, and they are about different things. total and
total_relation say how many matched and how far to trust the number. partial and
coverage say whether the index was fully read. A response can be exact and partial, or approximate
and complete. A geographic or amount filter with more candidates than its scan checks makes the total approximate,
and a derived filter that part of the index cannot yet apply makes the answer partial, with
coverage: field_unavailable.
Article
Under view=full, the default. view=compact keeps the fields listed above.
| Field | Type | Description |
|---|---|---|
| activity_density | number | Density of active language, as extracted. |
| authors | string[] | Bylines as extracted; pass one back as the `author` filter. |
| cities | string[] | Mentioned cities; pass one back as the `city` filter. |
| collected_at | string | When the index saw the article, 14-digit YYYYMMDDHHMMSS. |
| countries | string[] | Mentioned countries, canonical two-character identifiers; pass one back as the `country` filter. |
| event_types | string[] | Event classes (events-v1) implied by the article's GKG themes; pass one back as the `event_type` filter. Precise rather than complete: an article can describe an event without carrying its class. |
| id | string | Stable article id; pass to /doc/{id} and /similar/{id}. |
| image | string | Lead image URL, when the publisher declared one. |
| language | string | Source language, ISO-639-3. |
| locations | string[] | Mentioned place names; pass one back as the `location` filter. /doc/{id} has the structured form. |
| matched_locations | MatchedLocation[] | Present when near or bbox filtered the request: the mentioned cities that matched, nearest first. |
| mentioned_dates | string[] | Dates referred to in the text (YYYY, YYYY-MM, YYYY-MM-DD or --MM-DD); pass one back as the `mentioned_date` filter. |
| negative_score | number | Share of negative language, as extracted. |
| organizations | string[] | Organization entities as extracted; pass one back as the `organization` filter. |
| persons | string[] | Person entities as extracted; pass one back as the `person` filter. |
| polarity | number | Emotional charge regardless of direction, as extracted. |
| positive_score | number | Share of positive language, as extracted. |
| published_at | string | 14-digit YYYYMMDDHHMMSS, UTC. |
| publisher_country | string | The publisher's canonical two-character country identifier. |
| regions | string[] | Mentioned regions (states, provinces); pass one back as the `region` filter. |
| score | number | Ranking score of this hit: BM25 for keyword, fused similarity for semantic, reciprocal-rank fusion for hybrid. Comparable within one response only; absent on a recency browse. |
| self_reference_density | number | Density of self-referential language, as extracted. |
| signals | Signals | Versioned signal intensities. Absent when the article predates the signal vocabulary. |
| source | string | Publisher domain, e.g. bbc.co.uk. |
| source_type | string | web | citation | academic_archive | defense_archive | journal_archive | non_textual | other. |
| story_id | string | The cluster this article belongs to. Pass to /stories/{id} to expand the story, or back as the `story_id` filter. |
| title | string | Headline as published. |
| tone | number | Overall tone, negative to positive, as extracted. |
| topics | string[] | Canonical topic identifiers; pass one back as the `topic` filter. |
| url | string | Canonical article URL. Link this; the publisher gets the visit. |
| word_count | integer | Words in the article body, when known. |
Every hit carries event_types, the event classes its themes
imply. matched_locations appears only when near or bbox filtered the
request: the cities that matched, nearest first under near, each with its distance.
Geographic search has the details.
Article bodies are never returned — url is where the article lives. For the long tail of an article's
metadata (quotations, structured locations, amounts, alternate URLs), fetch it by id with
/doc/{id}.
Example
curl -s -G "https://api.unzoi.com/search" \
-H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=port congestion" \
--data-urlencode "mode=hybrid" \
--data-urlencode "language=eng" \
--data-urlencode "from=2026-08-01" \
--data-urlencode "facets=source,country" \
--data-urlencode "limit=10" Errors
| Status | Meaning |
|---|---|
| 200 | Success |
| 400 | Invalid request, charged nothing. Bounds, enums, dates and offsets are validated BEFORE any index work, so this never means a partially-served query. `credit_budget_exceeded`: the call's worst case is more than `max_credits`. |
| 401 | Missing or invalid API key. `code` is `unauthorized`. |
| 402 | Account suspended (subscription canceled or payment failed). `code` is `suspended`. |
| 429 | Refused before running, and charged nothing. `credit_rate_limited`: this minute's credits are spent. `message_rate_limited`: too many free calls or refused requests this minute. `insufficient_credits`: the key stops at its monthly allocation and does not have this call's worst case left. `retry-after` says how long to wait. |
| 503 | The index could not be read for this request. This is NOT an empty result set and NOT a statement that nothing matched — retry rather than caching it as a negative. `coverage` says which part was unreadable. |
See Errors for how to back off correctly.