Skip to content
unzoi docs
Search and navigation
Start here
REST API
MCP
Limits and plans
Agent clients
SDKs
Guides

GET /doc/{id}

GET https://api.unzoi.com/doc/{id}

One article's full metadata.

Names, structured locations and dates, quotations, amounts, related media, links and alternate URLs. Equivalent to the `get_article` MCP tool. Credits: 1. Pass estimate=true to get the exact range without running the call, or max_credits to refuse anything costlier.

  • id — The article id, from any search result.

Search responses carry a bounded subset of an article's metadata, so a page of results stays a reasonable size even when one article contains fifty quotations. This endpoint returns everything.

The id is whatever a search result gave you. It begins with the article's collection timestamp (YYYYMMDDHHMMSS), which is how the sharded index routes a lookup to exactly one shard — so this is a cheap request regardless of how large the archive is.

Response

Everything an Article carries, plus:

Field Type Description
activity_density number Density of active language, as extracted.
authors string[] Bylines as extracted; pass one back as the `author` filter.
cities string[] Mentioned cities; pass one back as the `city` filter.
collected_at string When the index saw the article, 14-digit YYYYMMDDHHMMSS.
countries string[] Mentioned countries, canonical two-character identifiers; pass one back as the `country` filter.
event_types string[] Event classes (events-v1) implied by the article's GKG themes; pass one back as the `event_type` filter. Precise rather than complete: an article can describe an event without carrying its class.
id string Stable article id; pass to /doc/{id} and /similar/{id}.
image string Lead image URL, when the publisher declared one.
language string Source language, ISO-639-3.
locations string[] Mentioned place names; pass one back as the `location` filter. /doc/{id} has the structured form.
matched_locations MatchedLocation[] Present when near or bbox filtered the request: the mentioned cities that matched, nearest first.
mentioned_dates string[] Dates referred to in the text (YYYY, YYYY-MM, YYYY-MM-DD or --MM-DD); pass one back as the `mentioned_date` filter.
negative_score number Share of negative language, as extracted.
organizations string[] Organization entities as extracted; pass one back as the `organization` filter.
persons string[] Person entities as extracted; pass one back as the `person` filter.
polarity number Emotional charge regardless of direction, as extracted.
positive_score number Share of positive language, as extracted.
published_at string 14-digit YYYYMMDDHHMMSS, UTC.
publisher_country string The publisher's canonical two-character country identifier.
regions string[] Mentioned regions (states, provinces); pass one back as the `region` filter.
score number Ranking score of this hit: BM25 for keyword, fused similarity for semantic, reciprocal-rank fusion for hybrid. Comparable within one response only; absent on a recency browse.
self_reference_density number Density of self-referential language, as extracted.
signals Signals Versioned signal intensities. Absent when the article predates the signal vocabulary.
source string Publisher domain, e.g. bbc.co.uk.
source_type string web | citation | academic_archive | defense_archive | journal_archive | non_textual | other.
story_id string The cluster this article belongs to. Pass to /stories/{id} to expand the story, or back as the `story_id` filter.
title string Headline as published.
tone number Overall tone, negative to positive, as extracted.
topics string[] Canonical topic identifiers; pass one back as the `topic` filter.
url string Canonical article URL. Link this; the publisher gets the visit.
word_count integer Words in the article body, when known.
amounts Amount[] Numeric amounts with the object they describe.
amp_url string The AMP version of the article, when declared.
date_mentions DateMention[] Dates referred to in the text, with precision. Distinct from `published_at`.
links string[] Outbound links from the article body.
location_details ArticleLocation[] Every mentioned place with its kind, country, identifiers and coordinates.
mobile_url string The mobile version of the article, when declared.
names string[] Proper names as extracted; pass one back as the `name` filter.
quotes Quotation[] Quotations as extracted, with the verb that introduced them.
related_images string[] Further image URLs from the article.
social_images string[] Images declared for social sharing.
social_videos string[] Videos declared for social sharing.

What the extra fields are for

  • quotes — extracted quotations with the verb that introduced them, each {text, verb}. Useful for "who said what" without fetching and parsing the article yourself, and filterable in search via quote_verb.
  • location_details — the same locations as the search response, but structured: each {country, kind, latitude, location_id, longitude, name, region_id, subregion_id}, with the kind (country, region, city), coordinates, and a canonical identifier you can filter on with location_id.
  • date_mentions — dates mentioned in the text, each {date, precision}, where precision says how much of the date the text gave. Distinct from published_at: an article published today can be about a date next year.
  • amounts — numeric amounts with the object they describe ("20000 combat soldiers"), each {number, object, value}. value is the number as written, a string that keeps the source's precision. number is the same amount parsed, or null when it does not parse. Filterable by what it counts with amount_object, and by number with amount_min and amount_max, where one amount must satisfy every bound.
  • amp_url, mobile_url, links — alternate and outbound URLs, for a fetcher that needs a lighter page or wants to follow the article's own citations.

Example

curl -s "https://api.unzoi.com/doc/20260826073000-a1b2c3" -H "x-api-key: $UNZOI_KEY"

An id that does not exist returns 404 with code: not_found — a statement of absence, safe to cache, unlike a 503. Article bodies are never returned by any endpoint — the index holds metadata and derived signals, and url is where the text lives.