Entity and signal filtering
Two ways to narrow a query that do not make ranking worse: exact filters on extracted entities, and ranges over derived signals.
Filters beat more words
Every word you add to q dilutes ranking — the engine now has to satisfy more terms, and articles that
match your real intent but use one word differently fall away. A filter does not compete with ranking at all: it
removes non-matching documents before ranking runs.
# Diluted.
q=asml export licence netherlands august 2026
# Same intent, ranked cleanly.
q=export licence
&entity_id=organization:asml
&publisher_country=NL
&from=2026-08-01 Entity filters are exact
organization, person, country, location, topic and
the rest match the index's own extracted values, not free text. organization=asml works;
organization=ASML Holding NV probably does not. An entity id, entity_id=organization:asml,
matches every spelling the index groups under it.
The reliable way to learn the spelling is to ask for it, starting from the name you have:
curl -s -G "https://api.unzoi.com/entities" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "type=organization" \
--data-urlencode "q=ASML Holding NV" \
--data-urlencode "from=2026-08-01" | jq '.entities[] | {entity_id, canonical, aliases, confidence}' {
"entity_id": "organization:asml",
"canonical": "asml",
"aliases": [
{ "value": "asml", "count": 1240, "match": "exact" },
{ "value": "asml holding nv", "count": 31, "match": "exact" }
],
"confidence": 1.0
} /entities normalises the name and every value the index holds the same
way, groups the values that match under one entity_id, and says which spelling is most common. Several
groups means the name is ambiguous; an empty list means the index holds no such spelling in the window, which is not
the same as the entity being absent from the news.
Filter on the id
Pass the id, and the filter matches every spelling in the group at once:
curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=export licence" \
--data-urlencode "entity_id=organization:asml" \
--data-urlencode "from=2026-08-01" | jq '{resolved, credits_charged}' {
"resolved": [
{ "entity_id": "organization:asml", "aliases_used": ["asml", "asml holding nv"] }
],
"credits_charged": 2
} resolved is how you know the id matched. An empty aliases_used means the index holds no
spelling of it in your window, so the empty result says nothing about the news. Up to five ids go in one
entity_id, comma separated, and all of them must match. Each one costs the index query that resolves
it, which is why this call reports 2 credits rather than 1.
The raw value is still a filter, and still the fallback. organization=asml matches that one spelling
and costs nothing extra. Use it when you took the value off a result or a facet, when you mean exactly that spelling,
or when one entity is split across two ids and you are merging the two queries yourself.
GET /entities explains when that happens.
Facets are the second way, and they answer a different question — not "how does the index spell this?" but "who is in this coverage at all?":
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=semiconductor supply" \
--data-urlencode "facets=organization,person,country,topic" \
--data-urlencode "facet_limit=15" \
--data-urlencode "limit=1" | jq '.facets | map_values(.values)' Facets count over the whole match set, so this tells you what actually exists in the coverage before you filter on it — but only among the articles your query text matched. Resolve when you have a name; facet when you have a topic.
Locations: name or identifier
location, city and region match the extracted string
("Austin, Texas, United States"), which is readable and brittle. location_id matches a canonical
numeric identifier, which is neither — take it from an article's location_details via
/doc/{id} and use it for anything long-lived.
Signals
Every article carries derived intensities across three families — industry, business context, risk context. They are not sentiment and not a classification: they measure how strongly an article's language sits in a given register.
That makes them the tool for a question about character rather than about subject: "supply-chain news that reads as disruption", "earnings coverage that reads as uncertainty". Those are hard to write as keywords and easy to write as a range.
# Shipping coverage that reads as disruption, not as routine logistics.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=shipping" \
--data-urlencode "signals=supply_disruption:1.5:,transportation:0.5:" Several constraints are ANDed. Over MCP the same thing is an array, which a model emits more reliably:
{
"name": "search_news",
"arguments": {
"q": "shipping",
"signals": [
{ "name": "supply_disruption", "min": 1.5 },
{ "name": "transportation", "min": 0.5 }
]
}
} Asking for the top slice instead of guessing a threshold
There is no absolute scale to memorise. Intensities are relative, and the useful cut depends on the slice you are looking at: an intensity that is extreme for earnings coverage can be routine for coverage of a war. So ask for a percentile, and let the index find the number:
# The 5% of this coverage that reads most like financial uncertainty.
curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=quarterly results" \
--data-urlencode "from=2026-08-01" \
--data-urlencode "signal=financial_uncertainty" \
--data-urlencode "signal_percentile_min=95" \
--data-urlencode "limit=20" | jq '.signal_threshold' {
"name": "financial_uncertainty",
"percentile": 95,
"value": 2.31,
"approximate": true
}
The index works out the intensity at that percentile over everything the query and filters matched, then filters
with signal_min at that value. signal_threshold reports the value it used, so you can log
it and watch it move. approximate: true means several index partitions answered and the value is the
count-weighted mean of theirs rather than one exact percentile. signal_percentile_min takes 50 to
99.9, needs signal, and does not combine with signal_min. It works on
/search, /stories and /top-headlines, and on the equivalent tools.
A percentile is relative to the match set, and that is the point: the sharpest 5% of Maersk coverage and the sharpest 5% of all shipping coverage are different cuts, each right for its question. It also means the same request next month can resolve to a different value. When you need a cut that stays put, use an absolute number.
If you want an absolute number
Find it empirically, then pin it with signal_min:
# keyword, so total is an exact count rather than a candidate count
for min in 0.5 1.0 1.5 2.0 3.0; do
total=$(curl -s -G "https://api.unzoi.com/search" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=quarterly results" \
--data-urlencode "mode=keyword" \
--data-urlencode "from=2026-08-01" \
--data-urlencode "signal=financial_uncertainty" \
--data-urlencode "signal_min=$min" \
--data-urlencode "limit=1" | jq -r '.total')
echo "min=$min $total articles"
done
Pick the point where the count starts dropping sharply, then read the top few titles at that threshold to check they
are what you meant. Five requests, and you have a number you can defend. The value a percentile
request reported is a sensible place to centre the scan.
Events, amounts and places
Three more filters narrow by what the index derives from each article's own records rather than by an extracted
name. event_type is what happened, such as a bankruptcy, an outage or a
labor_strike, taken from the article's themes. amount_min and amount_max
bound a number the article mentions, and with amount_object they say what it counts.
near and bbox match the cities it mentions.
q=strike
&event_type=labor_strike # what happened
&near=51.9244,4.4777&radius_km=50 # near Rotterdam
&amount_object=workers&amount_min=1000 # at least 1,000 workers
&from=2026-09-01
Event classes are precise rather than complete, so keep q alongside when recall matters.
Event classes lists them and the themes each needs,
Amount bounds has the rule that one amount must satisfy every bound, and
Geographic search covers areas. An area or an amount bound is checked article by article,
which can add credits.
Combining them
The three narrowing tools are independent, and stacking them is usually better than deepening any one of them:
q=port congestion # what it is about
&mode=hybrid # how to rank it
&entity_id=organization:maersk # who it names
&publisher_country=SG # who published it
&language=eng # what language
&from=2026-08-01 # when
&signals=supply_disruption:1.5: # what register
Every one of those except q and mode is free, in the sense that it costs no relevance.
The whole thing is one call: 3 credits for the hybrid search, plus 1 or 2 for resolving the
entity_id, all reported in credits_charged.