Filters and signals
Filters narrow a query without making ranking fuzzier. Adding words to q costs relevance; adding a
filter costs nothing.
Every filter
These work on /search,
/stories,
/top-headlines and
/aggregate, and as arguments to the equivalent
MCP tools, with near, radius_km, bbox, amount_min and amount_max not on /aggregate.
| Name | Type | Description |
|---|---|---|
| source | string | Source domain, e.g. bbc.co.uk. |
| source_type | `web` · `citation` · `academic_archive` · `defense_archive` · `journal_archive` · `non_textual` · `other` | The kind of source: web | citation | academic_archive | defense_archive | journal_archive | non_textual | other. |
| publisher_country | string | The publisher's canonical two-character country identifier. |
| language | string | Source language, ISO-639-3, e.g. eng, fra, ara, zho. |
| story_id | string | A cluster id from /stories, to expand one story into its articles. |
| topic | string | An exact canonical topic identifier returned by the API. |
| organization | string | An exact organization entity. |
| person | string | An exact person entity. |
| country | string | A mentioned country's canonical two-character identifier. |
| author | string | An exact article author. |
| location | string | An exact mentioned location. |
| location_id | string | An exact canonical location identifier. |
| city | string | An exact mentioned city. |
| region | string | An exact mentioned region. |
| name | string | An exact proper name. |
| mentioned_date | string | An exact date mentioned in the article text. |
| quote_verb | string | An exact verb introducing a quotation. |
| amount_object | string | An exact object described by a numeric amount. |
| near | string | A point, `latitude,longitude` in decimal degrees, e.g. 48.8566,2.3522. With radius_km: articles mentioning a city within that distance. On article results (/search, /top-headlines) each hit reports the cities that matched as matched_locations. |
| radius_km | number | With near: the distance in kilometres, 0.1-2000. |
| bbox | string | A box, `south,west,north,east` in decimal degrees: articles mentioning a city inside it. West greater than east crosses the antimeridian. Not with near. |
| event_type | `bankruptcy` · `earnings_report` · `ipo` · `cyber_attack` · `outage` · `industrial_accident` · `sanctions` · `trade_dispute` · `boycott` · `nationalization` · `privatization` · `money_laundering` · `executive_change` · `factory_closure` · `supply_shortage` · `labor_strike` · `antitrust_action` | An event class derived from the article's GKG themes (events-v1). Precise rather than complete: an article can describe an event without carrying its class. |
| amount_min | number | Articles mentioning a numeric amount of at least this value. With amount_max, one amount must satisfy both; with amount_object, that amount must describe the object. |
| amount_max | number | Articles mentioning a numeric amount of at most this value. With amount_min, one amount must satisfy both; with amount_object, that amount must describe the object. |
| estimate | boolean | Return this call's price instead of running it: the credit range, a breakdown, your remaining credits and whether the call would run. A preview costs 0 credits. |
| max_credits | integer | Refuse the call, at no cost, when its worst case would cost more than this many credits. The refusal carries the estimate. |
A parameter name that is not in this table is refused: 400 with code: unknown_parameter
and the name you probably meant. Values are checked where they can be — source_type,
event_type and signal against their enums, radius_km against its range,
near and bbox as coordinates — and matched exactly where they cannot.
Most of these match an attribute extracted from the article.
near, radius_km and bbox match
the cities an article mentions; event_type an
event class its themes imply;
amount_min and amount_max the amounts it mentions. Those three are derived from records every article already carries, and each index gains them
when it is rebuilt. Until then, the part of the index built without them answers
coverage: field_unavailable.
Entity ids
Each entity filter above matches one spelling. entity_id matches them all: pass an id from
/entities, such as organization:asml, and it stands for every
spelling in that id's alias group. It works where the filters do, on /search, /stories,
/top-headlines and /aggregate, and on their tools.
| Name | Type | Description |
|---|---|---|
| entity_id | string | Up to five entity ids from resolve_entity, comma separated (over REST the parameter may also repeat). Each matches any spelling in its alias group in the requested window; several ids must all match. The spellings used are reported in `resolved`. |
- Up to five ids, comma separated. Over REST the parameter may also repeat.
- Each id matches any spelling in its alias group within the request window. Several ids must all match.
-
The response's
resolvedlists the spellings each id used. An id with none in the window matches nothing, and itsaliases_usedis empty. -
Resolving an id costs a credit, sometimes two, which shows up in
credits_charged.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=export licence" \
--data-urlencode "entity_id=organization:asml" \
--data-urlencode "from=2026-08-01" | jq '.resolved' GET /entities covers how ids are formed, why two spellings of one name can still be two ids, and what to do about it. The raw filters stay the right tool for exactly one spelling, or a value read off a result.
Near a place
near with radius_km, or bbox, keeps articles that mention a city inside a
circle or a box, and each hit lists the cities that matched as matched_locations.
Geographic search covers the parameters, how the index matches them, why only cities count,
and what the check costs.
Event classes
event_type filters on what happened rather than on what an article is about. Every article carries
event_types, the classes its GDELT themes imply, and facets=event_type counts them over a
match set. It works on
/search, /stories, /top-headlines and /aggregate,
and on their tools. The filter takes one of
17 classes:
bankruptcy · earnings_report · ipo · cyber_attack · outage · industrial_accident · sanctions · trade_dispute · boycott · nationalization · privatization · money_laundering · executive_change · factory_closure · supply_shortage · labor_strike · antitrust_action
| Name | Type | Description |
|---|---|---|
| event_type | `bankruptcy` · `earnings_report` · `ipo` · `cyber_attack` · `outage` · `industrial_accident` · `sanctions` · `trade_dispute` · `boycott` · `nationalization` · `privatization` · `money_laundering` · `executive_change` · `factory_closure` · `supply_shortage` · `labor_strike` · `antitrust_action` | An event class derived from the article's GKG themes (events-v1). Precise rather than complete: an article can describe an event without carrying its class. |
How a class is derived
Classes come from each article's GKG themes, through a fixed table named events-v1. A class
needs one of its themes. Where a theme alone is ambiguous, the class also needs one theme from a second list: an
APPOINTMENT is an executive_change only when the article is also about a chief executive
or a managing director, because a minister's appointment is not one. Themes match whatever their case.
| Class | Needs one of | And one of |
|---|---|---|
bankruptcy | ECON_BANKRUPTCY | — |
earnings_report | ECON_EARNINGSREPORT | — |
ipo | ECON_IPO | — |
cyber_attack | CYBER_ATTACK | — |
outage | POWER_OUTAGE, INTERNET_BLACKOUT, PHONE_OUTAGE, MANMADE_DISASTER_POWER_OUTAGE, MANMADE_DISASTER_POWER_OUTAGES, MANMADE_DISASTER_POWER_BLACKOUT, MANMADE_DISASTER_POWER_DISRUPTION, MANMADE_DISASTER_DISRUPTION_OF_POWER, MANMADE_DISASTER_SERVICE_DISRUPTION | — |
industrial_accident | EMERG_INDUSTRIALACCIDENT, MANMADE_DISASTER_INDUSTRIAL_ACCIDENT, MANMADE_DISASTER_INDUSTRIAL_DISASTER, MANMADE_DISASTER_INDUSTRIAL_CATASTROPHE | — |
sanctions | SANCTIONS | — |
trade_dispute | ECON_TRADE_DISPUTE | — |
boycott | ECON_BOYCOTT | — |
nationalization | ECON_NATIONALIZE | — |
privatization | PRIVATIZATION | — |
money_laundering | ECON_MONEYLAUNDERING | — |
executive_change | RESIGNATION, APPOINTMENT | TAX_FNCACT_CEO, TAX_FNCACT_CHIEF_EXECUTIVE, TAX_FNCACT_CHIEF_EXECUTIVE_OFFICER, TAX_FNCACT_EXECUTIVE_OFFICER, TAX_FNCACT_MANAGING_DIRECTOR |
factory_closure | CLOSURE | WB_1281_MANUFACTURING, TAX_FNCACT_MANUFACTURER, TAX_FNCACT_FACTORY_WORKER, TAX_FNCACT_FACTORY_WORKERS |
supply_shortage | SHORTAGE | WB_1281_MANUFACTURING, TAX_FNCACT_MANUFACTURER, CRISISLEX_T06_SUPPLIES, WB_1353_PHARMACEUTICAL_SUPPLY_CHAIN, WB_194_AGRICULTURAL_SUPPLY_CHAIN, WB_2605_SUPPLY_CHAIN_ANALYSIS, WB_2628_AGRICULTURE_SUPPLY_CHAINS |
labor_strike | STRIKE | ECON_UNIONS, TAX_FNCACT_WORKERS, TAX_FNCACT_WORKER, TAX_FNCACT_EMPLOYEES |
antitrust_action | WB_2101_ANTITRUST, WB_1743_ANTICARTEL_ENFORCEMENT, WB_1031_COMPETITION_LAW, ECON_MONOPOLY | TAX_FNCACT_REGULATOR, TAX_FNCACT_REGULATORS |
Every theme in the table is spelled as it appears in GDELT's own theme vocabulary. The filter is exact against the
table: an article it returns carries the themes its class needs. A class that is not one of the
17 is refused with invalid_parameter and the closest class.
Some events are not classes. Acquisitions, funding rounds, layoffs, lawsuits, contract awards,
product launches and factory openings have no theme that names them reliably, so the table leaves them out rather
than guess. Search for those with q. The table is versioned, and a new version can add classes.
# Bankruptcies reported in English since August, one row per story.
curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "event_type=bankruptcy" \
--data-urlencode "language=eng" \
--data-urlencode "from=2026-08-01"
# Which kinds of event semiconductor coverage holds.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=semiconductor" \
--data-urlencode "from=2026-08-01" \
--data-urlencode "facets=event_type" \
--data-urlencode "limit=1" | jq '.facets.event_type.values' Amount bounds
amount_min and amount_max match the numbers articles mention: the amounts GDELT extracts
with what they count, such as "20000 combat soldiers".
| Name | Type | Description |
|---|---|---|
| amount_min | number | Articles mentioning a numeric amount of at least this value. With amount_max, one amount must satisfy both; with amount_object, that amount must describe the object. |
| amount_max | number | Articles mentioning a numeric amount of at most this value. With amount_min, one amount must satisfy both; with amount_object, that amount must describe the object. |
- One amount satisfies every bound.
amount_min=1000&amount_max=1500matches an article that mentions 1,200 of something. It does not match one that mentions 500 of one thing and 2,000 of another. - With
amount_object, that same amount describes the object.amount_object=workers&amount_min=1000needs an amount of at least 1,000 whose object isworkers. An article with 5,000 of something else and a smaller number of workers does not match.amount_objectis exact, whatever the case; take the value fromamounts[].objecton/doc/{id}. - The value compared is
amounts[].numberon/doc/{id}, the amount parsed as a number. An amount that does not parse, whosenumberisnull, never satisfies a bound. - A number has no unit but its object.
amount_min=1000000alone matches a million dollars, a million barrels and a million people alike, so pair it withamount_object. amount_minaboveamount_maxis refused withinvalid_parameter.
The index narrows to articles whose smallest and largest amounts could satisfy the bounds, then checks every
candidate against its document. That is the same scan a geographic filter takes, with the same effect on
total_relation, facets and credits_charged:
Exactness and cost.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "amount_object=workers" \
--data-urlencode "amount_min=1000" \
--data-urlencode "from=2026-09-01" | jq '{total, total_relation, credits_charged}' Where the values come from
| Filter | Vocabulary |
|---|---|
language | ISO 639-3, three letters: eng, fra, ara, zho. Not en. |
country, publisher_country | Two-character country identifiers. country is a country mentioned in the article; publisher_country is where the outlet is. |
topic | Canonical topic identifiers, returned in every article's topics. Coarse and machine-generated — treat them as a filter, not as a taxonomy to show a user. |
organization, person, name | Extracted entities, lower-cased. Read them off a result, or resolve a name with /entities, which groups the spellings the index holds under one entity id and says which is most common. |
location, city, region | Full location strings as extracted ("Austin, Texas, United States"). For stable matching prefer location_id. |
location_id | A numeric canonical identifier, from an article's location_details via /doc/{id}. |
source | The publisher's domain, e.g. bbc.co.uk. No scheme, no www.. |
story_id | A cluster id from /stories or from any article's story_id. Also a facet field, and the key /stories/{id} expands. |
mentioned_date | A date mentioned in the text, YYYY-MM-DD — not the publication date. |
quote_verb, amount_object | From an article's quotes and amounts; see /doc/{id}. |
near, bbox | Decimal degrees: latitude,longitude, or south,west,north,east. See Geographic search. |
event_type | One of the event classes. An article's event_types lists the ones it carries. |
amount_min, amount_max | Plain numbers, compared with amounts[].number. See Amount bounds. |
Signals
Every article carries derived intensities across three families. They are not sentiment and not a classification — they measure how strongly an article's language sits in a given register, so range-filtering on them finds coverage with a particular character rather than a particular word.
Each article reports the version that produced them (signals.version), so a change to the model is
visible in the data rather than silently shifting your thresholds.
Industry
What sector the coverage is about.
economy · healthcare · agriculture · labor · environment · energy · transportation · real_estate · finance · defense · science_technology · trade · public_sector
Business context
The financial register the article is written in.
financial_uncertainty · financial_negative · financial_positive · legal_litigation · financial_stability_stress
Risk context
Disruption, threat and instability language.
anxiety · conflict · supply_disruption · cyber_incident · climate · health_security · governance_risk
Filtering on them
| Name | Type | Description |
|---|---|---|
| signal | `economy` · `healthcare` · `agriculture` · `labor` · `environment` · `energy` · `transportation` · `real_estate` · `finance` · `defense` · `science_technology` · `trade` · `public_sector` · `financial_uncertainty` · `financial_negative` · `financial_positive` · `legal_litigation` · `financial_stability_stress` · `anxiety` · `conflict` · `supply_disruption` · `cyber_incident` · `climate` · `health_security` · `governance_risk` | A normalized business signal to range-filter on, e.g. finance, energy, conflict. An unknown name is rejected. |
| signal_min | number | Inclusive lower bound for `signal`. |
| signal_max | number | Inclusive upper bound for `signal`. |
| signals | string | Several signal ranges at once, ANDed: `name[:min[:max]]`, comma separated, e.g. `finance:1.5:,energy::2`. |
| signal_percentile_min | number | With `signal`: keep only articles at or above this percentile of the signal's intensity over the match set (50-99.9). The threshold used is reported as `signal_threshold`. Not with `signal_min`. |
A name that is not one of the 25 above is refused with code: unknown_signal and the
closest published name; signal_min or signal_max without signal is
invalid_parameter. Neither is matched to nothing any more.
One constraint:
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=quarterly results" \
--data-urlencode "signal=financial_uncertainty" \
--data-urlencode "signal_min=2.0" Several at once, ANDed, as name[:min[:max]] separated by commas — an empty bound means unbounded:
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=shipping" \
--data-urlencode "signals=supply_disruption:1.5:,transportation:0.5:3" Over MCP the same thing is an array of objects, which is easier for a model to emit correctly:
{
"name": "search_news",
"arguments": {
"q": "shipping",
"signals": [
{ "name": "supply_disruption", "min": 1.5 },
{ "name": "transportation", "min": 0.5, "max": 3 }
]
}
} Relative thresholds
There is no absolute scale to memorise — intensities are relative, and the useful thresholds depend on the corpus
slice you are looking at. signal_percentile_min asks for the slice instead. With
signal, it keeps only articles at or above that percentile of the signal's intensity over the match
set, and the response reports what that came to as signal_threshold:
curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=shipping" \
--data-urlencode "from=2026-08-01" \
--data-urlencode "signal=supply_disruption" \
--data-urlencode "signal_percentile_min=95" | Field | Type | Description |
|---|---|---|
| approximate | boolean | The percentile was a count-weighted mean across indexes rather than one exact percentile. |
| name | string | The signal. |
| percentile | number | The percentile requested. |
| value | number | The intensity at that percentile over the match set; the page was filtered with `signal_min` at this value. Null when no matching article carries the signal, in which case nothing matches. |
It takes 50 to 99.9, on /search, /stories and /top-headlines. Without
signal, or alongside signal_min, it is invalid_parameter.
Entity and signal filtering covers when a percentile
is the right cut, and how to find an absolute number when it is not.