GET /doc/{id}
One article's full metadata.
Names, structured locations and dates, quotations, amounts, related media, links and alternate URLs. Equivalent to the `get_article` MCP tool. Credits: 1. Pass estimate=true to get the exact range without running the call, or max_credits to refuse anything costlier.
-
id— The article id, from any search result.
Search responses carry a bounded subset of an article's metadata, so a page of results stays a reasonable size even when one article contains fifty quotations. This endpoint returns everything.
The id is whatever a search result gave you. It begins with the article's collection timestamp
(YYYYMMDDHHMMSS), which is how the sharded index routes a lookup to exactly one shard — so this is a
cheap request regardless of how large the archive is.
Response
Everything an Article carries, plus:
| Field | Type | Description |
|---|---|---|
| activity_density | number | Density of active language, as extracted. |
| authors | string[] | Bylines as extracted; pass one back as the `author` filter. |
| cities | string[] | Mentioned cities; pass one back as the `city` filter. |
| collected_at | string | When the index saw the article, 14-digit YYYYMMDDHHMMSS. |
| countries | string[] | Mentioned countries, canonical two-character identifiers; pass one back as the `country` filter. |
| event_types | string[] | Event classes (events-v1) implied by the article's GKG themes; pass one back as the `event_type` filter. Precise rather than complete: an article can describe an event without carrying its class. |
| id | string | Stable article id; pass to /doc/{id} and /similar/{id}. |
| image | string | Lead image URL, when the publisher declared one. |
| language | string | Source language, ISO-639-3. |
| locations | string[] | Mentioned place names; pass one back as the `location` filter. /doc/{id} has the structured form. |
| matched_locations | MatchedLocation[] | Present when near or bbox filtered the request: the mentioned cities that matched, nearest first. |
| mentioned_dates | string[] | Dates referred to in the text (YYYY, YYYY-MM, YYYY-MM-DD or --MM-DD); pass one back as the `mentioned_date` filter. |
| negative_score | number | Share of negative language, as extracted. |
| organizations | string[] | Organization entities as extracted; pass one back as the `organization` filter. |
| persons | string[] | Person entities as extracted; pass one back as the `person` filter. |
| polarity | number | Emotional charge regardless of direction, as extracted. |
| positive_score | number | Share of positive language, as extracted. |
| published_at | string | 14-digit YYYYMMDDHHMMSS, UTC. |
| publisher_country | string | The publisher's canonical two-character country identifier. |
| regions | string[] | Mentioned regions (states, provinces); pass one back as the `region` filter. |
| score | number | Ranking score of this hit: BM25 for keyword, fused similarity for semantic, reciprocal-rank fusion for hybrid. Comparable within one response only; absent on a recency browse. |
| self_reference_density | number | Density of self-referential language, as extracted. |
| signals | Signals | Versioned signal intensities. Absent when the article predates the signal vocabulary. |
| source | string | Publisher domain, e.g. bbc.co.uk. |
| source_type | string | web | citation | academic_archive | defense_archive | journal_archive | non_textual | other. |
| story_id | string | The cluster this article belongs to. Pass to /stories/{id} to expand the story, or back as the `story_id` filter. |
| title | string | Headline as published. |
| tone | number | Overall tone, negative to positive, as extracted. |
| topics | string[] | Canonical topic identifiers; pass one back as the `topic` filter. |
| url | string | Canonical article URL. Link this; the publisher gets the visit. |
| word_count | integer | Words in the article body, when known. |
| amounts | Amount[] | Numeric amounts with the object they describe. |
| amp_url | string | The AMP version of the article, when declared. |
| date_mentions | DateMention[] | Dates referred to in the text, with precision. Distinct from `published_at`. |
| links | string[] | Outbound links from the article body. |
| location_details | ArticleLocation[] | Every mentioned place with its kind, country, identifiers and coordinates. |
| mobile_url | string | The mobile version of the article, when declared. |
| names | string[] | Proper names as extracted; pass one back as the `name` filter. |
| quotes | Quotation[] | Quotations as extracted, with the verb that introduced them. |
| related_images | string[] | Further image URLs from the article. |
| social_images | string[] | Images declared for social sharing. |
| social_videos | string[] | Videos declared for social sharing. |
What the extra fields are for
-
quotes— extracted quotations with the verb that introduced them, each{text, verb}. Useful for "who said what" without fetching and parsing the article yourself, and filterable in search viaquote_verb. -
location_details— the same locations as the search response, but structured: each{country, kind, latitude, location_id, longitude, name, region_id, subregion_id}, with the kind (country, region, city), coordinates, and a canonical identifier you can filter on withlocation_id. -
date_mentions— dates mentioned in the text, each{date, precision}, whereprecisionsays how much of the date the text gave. Distinct frompublished_at: an article published today can be about a date next year. -
amounts— numeric amounts with the object they describe ("20000 combat soldiers"), each{number, object, value}.valueis the number as written, a string that keeps the source's precision.numberis the same amount parsed, ornullwhen it does not parse. Filterable by what it counts withamount_object, and bynumberwithamount_minandamount_max, where one amount must satisfy every bound. -
amp_url,mobile_url,links— alternate and outbound URLs, for a fetcher that needs a lighter page or wants to follow the article's own citations.
Example
curl -s "https://api.unzoi.com/doc/20260826073000-a1b2c3" -H "x-api-key: $UNZOI_KEY"
An id that does not exist returns 404 with code: not_found — a statement of absence, safe
to cache, unlike a 503. Article bodies are never returned by any endpoint — the index holds metadata
and derived signals, and url is where the text lives.