Skip to content
unzoi docs
Search and navigation
Start here
REST API
MCP
Limits and plans
Agent clients
SDKs
Guides

GET /search

GET https://api.unzoi.com/search

Search articles.

Ranked articles with topics, entities, locations and signals. Equivalent to the `search_news` MCP tool. Credits: 1 for keyword ranking or a browse, 2 for semantic, 3 for hybrid; 1-4 with near, bbox, amount_min or amount_max, whose candidates are checked in batches of 500; a cursor page costs 1-2 to start and 1-3 after; +1-2 per entity_id; +1 when signal_percentile_min resolves a threshold; +1 per full 1000 rows of offset. Pass estimate=true to get the exact range without running the call, or max_credits to refuse anything costlier.

Ranking

Name Type Description
q string Query text. Omit for a recency browse.
mode `keyword` · `semantic` · `hybrid` · `recent` Ranking mode. `hybrid` is usually right when a model wrote the query, because models paraphrase. `recent` orders newest-first and is implied when `q` is omitted.

mode is the parameter worth thinking about. Use hybrid when a model wrote the query — models paraphrase, and keyword ranking cannot bridge "central bank tightening" to an article that says "Fed raises rates". Use keyword when a human gave you an exact phrase, a product name or a proper noun they want matched literally. The guide goes deeper.

Omitting q entirely turns this into a recency browse — mode is reported as recent, and you can ask for it by name — which is what /top-headlines does explicitly. The same is true of the search_news tool: q is optional there too, so search_news { story_id: … } lists a story's members with no query at all.

Filters

Exact matches on indexed attributes. Adding words to q makes ranking fuzzier; adding a filter does not, and costs nothing.

Name Type Description
source string Source domain, e.g. bbc.co.uk.
source_type `web` · `citation` · `academic_archive` · `defense_archive` · `journal_archive` · `non_textual` · `other` The kind of source: web | citation | academic_archive | defense_archive | journal_archive | non_textual | other.
publisher_country string The publisher's canonical two-character country identifier.
language string Source language, ISO-639-3, e.g. eng, fra, ara, zho.
story_id string A cluster id from /stories, to expand one story into its articles.
topic string An exact canonical topic identifier returned by the API.
organization string An exact organization entity.
person string An exact person entity.
country string A mentioned country's canonical two-character identifier.
author string An exact article author.
location string An exact mentioned location.
location_id string An exact canonical location identifier.
city string An exact mentioned city.
region string An exact mentioned region.
name string An exact proper name.
mentioned_date string An exact date mentioned in the article text.
quote_verb string An exact verb introducing a quotation.
amount_object string An exact object described by a numeric amount.
estimate boolean Return this call's price instead of running it: the credit range, a breakdown, your remaining credits and whether the call would run. A preview costs 0 credits.
max_credits integer Refuse the call, at no cost, when its worst case would cost more than this many credits. The refusal carries the estimate.

Values are canonical identifiers, not free text — take them from a previous response, resolve a name with /entities, or read them off a facet count. Filters and signals explains the vocabularies. A parameter name that is not in these tables is refused with 400 and code: unknown_parameter, not ignored.

Entity ids

Name Type Description
entity_id string Up to five entity ids from resolve_entity, comma separated (over REST the parameter may also repeat). Each matches any spelling in its alias group in the requested window; several ids must all match. The spellings used are reported in `resolved`.

An id from /entities matches every spelling in its alias group, where a filter matches one. The response's resolved lists the spellings each id used, and resolving them is counted in credits_charged. Filters and signals has the rules.

Place, event classes and amounts

Name Type Description
near string A point, `latitude,longitude` in decimal degrees, e.g. 48.8566,2.3522. With radius_km: articles mentioning a city within that distance. On article results (/search, /top-headlines) each hit reports the cities that matched as matched_locations.
radius_km number With near: the distance in kilometres, 0.1-2000.
bbox string A box, `south,west,north,east` in decimal degrees: articles mentioning a city inside it. West greater than east crosses the antimeridian. Not with near.
event_type `bankruptcy` · `earnings_report` · `ipo` · `cyber_attack` · `outage` · `industrial_accident` · `sanctions` · `trade_dispute` · `boycott` · `nationalization` · `privatization` · `money_laundering` · `executive_change` · `factory_closure` · `supply_shortage` · `labor_strike` · `antitrust_action` An event class derived from the article's GKG themes (events-v1). Precise rather than complete: an article can describe an event without carrying its class.
amount_min number Articles mentioning a numeric amount of at least this value. With amount_max, one amount must satisfy both; with amount_object, that amount must describe the object.
amount_max number Articles mentioning a numeric amount of at most this value. With amount_min, one amount must satisfy both; with amount_object, that amount must describe the object.

These filter on fields the index derives from each article's own records: the cities it mentions, the event classes its themes imply, and the amounts it mentions. Geographic search covers near, radius_km and bbox; event classes and amount bounds are in Filters and signals. A geographic or amount filter is checked against each candidate article, which can make total_relation approximate and adds to credits_charged: Exactness and cost.

Signals

Every article carries versioned industry, business-context and risk-context intensities. Range-filtering on them finds coverage with a particular character rather than a particular word.

Name Type Description
signal `economy` · `healthcare` · `agriculture` · `labor` · `environment` · `energy` · `transportation` · `real_estate` · `finance` · `defense` · `science_technology` · `trade` · `public_sector` · `financial_uncertainty` · `financial_negative` · `financial_positive` · `legal_litigation` · `financial_stability_stress` · `anxiety` · `conflict` · `supply_disruption` · `cyber_incident` · `climate` · `health_security` · `governance_risk` A normalized business signal to range-filter on, e.g. finance, energy, conflict. An unknown name is rejected.
signal_min number Inclusive lower bound for `signal`.
signal_max number Inclusive upper bound for `signal`.
signals string Several signal ranges at once, ANDed: `name[:min[:max]]`, comma separated, e.g. `finance:1.5:,energy::2`.
signal_percentile_min number With `signal`: keep only articles at or above this percentile of the signal's intensity over the match set (50-99.9). The threshold used is reported as `signal_threshold`. Not with `signal_min`.

Use signal + signal_min/signal_max for one constraint, signals for several at once, ANDed, or signal + signal_percentile_min for the top slice of whatever matched, with the value it resolved to reported as signal_threshold. See Entity and signal filtering.

Time range

Name Type Description
from string Lower time bound, e.g. 2026-07-09 or 20260709120000. Clamped forward to what the caller's plan may reach; see `history_days` and `from_clamped` in the response.
to string Upper time bound, same formats as `from`.

Paging, shape and facets

Name Type Description
offset integer Zero-based result offset, up to 100000.
limit integer Results per page, 1-100 (default 10).
cursor string Keyset paging for a newest-first browse (no `q`): pass `start` for the first page, then the `next_cursor` the response carried. Pages are deterministic and disjoint even across articles sharing one timestamp. The first page costs two index queries and each later page three. Not with `offset`, `q`, or a ranked `mode`.
facets string Comma-separated fields to count over the full match set, up to 8.
facet_limit integer Values per facet, 1-100 (default 10).
view `full` · `compact` How much of each hit to return. `full` (default) is the whole Article; `compact` keeps id, title, url, source, published_at, language, story_id and score. Use compact when results feed a model.

offset pages any query. cursor pages a browse with no q, newest first: pass start, then each next_cursor, and the pages stay disjoint while new articles arrive. The two do not combine. Pagination and facets covers cursor paging, what the counts mean, and when they are approximate.

view changes the shape of a hit, not the match set. compact keeps id, language, published_at, score, source, story_id, title, url and drops the rest: the entity lists, the signal map, the tone scores. Use it when results feed a model that is deciding what to read next — the full record is one /doc/{id} away — and leave it at full when your code reads the fields.

Response

Field Type Description
attribution string Required source attribution for any use of these results.
coverage `complete` · `timed_out` · `range_too_wide` · `recent_unavailable` · `field_unavailable` Why the answer is or is not complete, named for what you can do about it. `complete`: everything in range was read (or provably could not have changed the page). `timed_out`: the query ran out of budget -- retry, or narrow the range. `range_too_wide`: the range spans more of the corpus than one request may read -- narrow it; retrying unchanged will not help. `recent_unavailable`: the most recent data in range could not be read -- retry shortly; widening the range will not help. `field_unavailable`: a filter needs a field part of the index was built without (geographic, event-class and amount fields arrive with each shard's rebuild) -- the answer covers only the rebuilt part; retrying before the rebuild will not help.
credits_charged integer Index queries this answer took, and what the call is charged (also the `x-credits-charged` header). One for a plain search; more with `entity_id`, `signal_percentile_min`, a cursor page, a geographic or amount filter (a query per batch of 500 candidates checked), or a derived-field filter on stories, semantic or hybrid search and aggregate (one query to check the index holds the field).
facets object Present only when `facets` was requested; keyed by the public field name.
from_clamped boolean Whether a `from` you supplied was moved forward to that boundary. Distinguishes a plan limit from a corpus with no coverage.
has_more boolean Whether another page exists. Stop paging when false.
history_days integer How far back the caller's plan may query; null = the full archive.
limit integer The page size that was applied.
mode `keyword` · `semantic` · `hybrid` · `recent` · `similar` The ranking that was applied. `recent` when no `q` was given; `similar` from /similar/{id}.
next_cursor string On a cursor-paged browse: pass back as `cursor` for the next page. Absent when there is no further page or `cursor` was not used.
offset integer The zero-based offset this page starts at.
partial boolean Whether this answer is incomplete. `true` means part of the index could not be read, so an empty or short result set is NOT evidence of absence -- retry rather than caching it as a negative.
resolved ResolvedEntity[] Present when `entity_id` was used: the spellings each id matched.
results Article[] This page of hits, best first.
signal_threshold SignalThreshold Present when `signal_percentile_min` was used: the threshold it resolved to.
total integer Matches in the whole set (see `total_relation`), not on this page.
total_relation `exact` · `approximate` `exact` for keyword and metadata queries, and for geographic and amount filters when every candidate was checked; `approximate` for ANN-derived modes, which rank a bounded candidate set, and for a checked filter whose candidates exceeded the scan. `/stories` is always `approximate`: stories are collapsed from a bounded ranked set of articles.

Two pairs of fields describe the answer, and they are about different things. total and total_relation say how many matched and how far to trust the number. partial and coverage say whether the index was fully read. A response can be exact and partial, or approximate and complete. A geographic or amount filter with more candidates than its scan checks makes the total approximate, and a derived filter that part of the index cannot yet apply makes the answer partial, with coverage: field_unavailable.

Article

Under view=full, the default. view=compact keeps the fields listed above.

Field Type Description
activity_density number Density of active language, as extracted.
authors string[] Bylines as extracted; pass one back as the `author` filter.
cities string[] Mentioned cities; pass one back as the `city` filter.
collected_at string When the index saw the article, 14-digit YYYYMMDDHHMMSS.
countries string[] Mentioned countries, canonical two-character identifiers; pass one back as the `country` filter.
event_types string[] Event classes (events-v1) implied by the article's GKG themes; pass one back as the `event_type` filter. Precise rather than complete: an article can describe an event without carrying its class.
id string Stable article id; pass to /doc/{id} and /similar/{id}.
image string Lead image URL, when the publisher declared one.
language string Source language, ISO-639-3.
locations string[] Mentioned place names; pass one back as the `location` filter. /doc/{id} has the structured form.
matched_locations MatchedLocation[] Present when near or bbox filtered the request: the mentioned cities that matched, nearest first.
mentioned_dates string[] Dates referred to in the text (YYYY, YYYY-MM, YYYY-MM-DD or --MM-DD); pass one back as the `mentioned_date` filter.
negative_score number Share of negative language, as extracted.
organizations string[] Organization entities as extracted; pass one back as the `organization` filter.
persons string[] Person entities as extracted; pass one back as the `person` filter.
polarity number Emotional charge regardless of direction, as extracted.
positive_score number Share of positive language, as extracted.
published_at string 14-digit YYYYMMDDHHMMSS, UTC.
publisher_country string The publisher's canonical two-character country identifier.
regions string[] Mentioned regions (states, provinces); pass one back as the `region` filter.
score number Ranking score of this hit: BM25 for keyword, fused similarity for semantic, reciprocal-rank fusion for hybrid. Comparable within one response only; absent on a recency browse.
self_reference_density number Density of self-referential language, as extracted.
signals Signals Versioned signal intensities. Absent when the article predates the signal vocabulary.
source string Publisher domain, e.g. bbc.co.uk.
source_type string web | citation | academic_archive | defense_archive | journal_archive | non_textual | other.
story_id string The cluster this article belongs to. Pass to /stories/{id} to expand the story, or back as the `story_id` filter.
title string Headline as published.
tone number Overall tone, negative to positive, as extracted.
topics string[] Canonical topic identifiers; pass one back as the `topic` filter.
url string Canonical article URL. Link this; the publisher gets the visit.
word_count integer Words in the article body, when known.

Every hit carries event_types, the event classes its themes imply. matched_locations appears only when near or bbox filtered the request: the cities that matched, nearest first under near, each with its distance. Geographic search has the details.

Article bodies are never returned — url is where the article lives. For the long tail of an article's metadata (quotations, structured locations, amounts, alternate URLs), fetch it by id with /doc/{id}.

Example

curl -s -G "https://api.unzoi.com/search" \
  -H "x-api-key: $UNZOI_KEY" \
  --data-urlencode "q=port congestion" \
  --data-urlencode "mode=hybrid" \
  --data-urlencode "language=eng" \
  --data-urlencode "from=2026-08-01" \
  --data-urlencode "facets=source,country" \
  --data-urlencode "limit=10"

Errors

Status Meaning
200 Success
400 Invalid request, charged nothing. Bounds, enums, dates and offsets are validated BEFORE any index work, so this never means a partially-served query. `credit_budget_exceeded`: the call's worst case is more than `max_credits`.
401 Missing or invalid API key. `code` is `unauthorized`.
402 Account suspended (subscription canceled or payment failed). `code` is `suspended`.
429 Refused before running, and charged nothing. `credit_rate_limited`: this minute's credits are spent. `message_rate_limited`: too many free calls or refused requests this minute. `insufficient_credits`: the key stops at its monthly allocation and does not have this call's worst case left. `retry-after` says how long to wait.
503 The index could not be read for this request. This is NOT an empty result set and NOT a statement that nothing matched — retry rather than caching it as a negative. `coverage` says which part was unreadable.

See Errors for how to back off correctly.