Monitoring a topic
You can poll, or you can let the index poll for you. Either way, five things make a monitor correct: a watermark, a deduplication key, a rule for when a story you already reported is worth mentioning again, a check that the poll was complete before the watermark moves, and an interval you can afford.
Create a watch instead
A watch is the loop on this page, run by the index against your key. It runs a
/stories query every interval_minutes, overlaps each window with the last, remembers
stories by story_id, reports a story again when it crosses 5, 15, 40 or 100 outlets or grows by ten,
and moves its watermark only on a complete run. Four of the five decisions are made for you. The interval is still
yours, and your plan sets its floor.
curl -s -X POST "https://api.unzoi.com/watches" -H "x-api-key: $UNZOI_KEY" \
-H "content-type: application/json" \
-d '{
"label": "port congestion",
"q": "port congestion OR container backlog",
"interval_minutes": 15,
"minimum_outlets": 3,
"webhook_url": "https://hooks.example.com/unzoi"
}'
The response carries the watch's id and, because it has a webhook_url, its
webhook_secret. The secret is returned once. Store it then. From an agent, the same request is the
create_watch tool.
Reading what it found
Every event lands in the watch's feed, webhook or not: new_story when a story first reaches
minimum_outlets, story_growth when it grows. Read it with
GET /watches/{id}/events, or the watch_events
tool.
curl -s "https://api.unzoi.com/watches/$WATCH_ID/events?limit=20" -H "x-api-key: $UNZOI_KEY" \
| jq -r '.events[] | "\(.kind)\t\(.story.outlets)\t\(.story.title)"'
Events are newest first and kept seven days. The feed is in the order events were recorded: each id begins with
when the run that recorded it was due. Read from the top until you reach an id you have already handled, and stop
there. Each page you read is a request. label names the watch in your lists; name in a
watch is the proper-name filter, as on /stories.
Verifying a webhook
With a webhook_url, each event is also POSTed to you. Check x-unzoi-signature against the
raw body before trusting it, answer 2xx within ten seconds, and deduplicate on
x-unzoi-delivery: a failed delivery is sent again after 1, 5 and 25 minutes.
import { createHmac, timingSafeEqual } from "node:crypto";
// rawBody is the request body exactly as received, before JSON.parse.
function verify(rawBody: Buffer, header: string | undefined, secret: string): boolean {
const expected = Buffer.from("sha256=" + createHmac("sha256", secret).update(rawBody).digest("hex"));
const received = Buffer.from(header ?? "");
return received.length === expected.length && timingSafeEqual(received, expected);
} Watches has the Python version, the payload, and what happens when deliveries keep failing.
When to poll instead
The loop below is the portable path. Keep it when:
-
you run the index yourself. Watches are served by the hosted API; the local stdio
server answers the watch tools with
unsupported. - the query needs something a watch does not carry: its own time window, a
sort,signal_percentile_min. - the growth rule should be yours rather than the watch's steps.
- one interval matches more than 100 stories. A watch's run reads the newest 100 in its window.
The loop
import { setTimeout as sleep } from "node:timers/promises";
const BASE = "https://api.unzoi.com";
const headers = { "x-api-key": process.env.UNZOI_KEY! };
// The watermark. Persist it — an in-memory one restarts your feed from scratch
// on every deploy, which is how a monitor floods a channel at 3am.
let watermark = loadWatermark() ?? stamp(Date.now() - 3600_000);
// What you have already told someone about, and how big it was then.
const known = new Map<string, { outlets: number; count: number; notifiedAt: number }>(loadKnown());
let stalePolls = 0;
async function poll() {
const url = new URL(BASE + "/stories");
url.searchParams.set("q", "port congestion OR container backlog");
url.searchParams.set("mode", "hybrid");
// Overlap deliberately: articles are indexed a little after they are
// published, so a watermark exactly at the last poll misses the stragglers.
url.searchParams.set("from", stamp(parse(watermark) - 15 * 60_000));
url.searchParams.set("limit", "100");
const response = await fetch(url, { headers });
if (response.status === 429) {
if ((await response.clone().json()).code === "insufficient_credits") {
console.error("out of credits — pausing until the month resets");
return;
}
await sleep(1000);
return;
}
if (!response.ok) return; // transient; the next tick catches up
const page = await response.json();
for (const story of page.stories) {
const before = known.get(story.story_id);
if (!before) {
// New event. Dedup on story_id, not on article id: forty outlets running
// one wire story is ONE thing that happened, and forty notifications is spam.
known.set(story.story_id, { outlets: story.outlets, count: story.count, notifiedAt: Date.now() });
notify(story, "new");
} else if (grew(before.outlets, story.outlets)) {
// Something you already reported got bigger. Expand it for the current
// picture — 2 credits, and limit=1 keeps the body small.
const detail = await (await fetch(`${BASE}/stories/${story.story_id}?limit=1`, { headers })).json();
known.set(story.story_id, { outlets: detail.outlets, count: detail.count, notifiedAt: Date.now() });
notify(detail, "growth");
}
}
// Advance only on a complete poll. A poll that reached part of the index is
// not a poll of that interval, and advancing on it opens a silent gap.
if (page.partial || page.coverage !== "complete") {
if (++stalePolls >= 4) alertOperator(`${stalePolls} incomplete polls in a row (${page.coverage})`);
return;
}
stalePolls = 0;
watermark = stamp(Date.now());
saveWatermark(watermark);
saveKnown(prune(known));
}
// A story is worth mentioning again when its outlet count crosses a step, or
// has grown by ten since you last mentioned it.
const STEPS = [5, 15, 40];
const grew = (before: number, now: number) =>
now - before >= 10 || STEPS.some((step) => before < step && now >= step);
// Bound the map: a story quiet for a day is over, and an unbounded map is a
// memory leak with a slow fuse.
const prune = (m: typeof known) =>
new Map([...m].filter(([, s]) => Date.now() - s.notifiedAt < 24 * 3600_000));
for (;;) {
await poll();
await sleep(15 * 60_000);
}
const stamp = (ms: number) => new Date(ms).toISOString().replace(/[-:T]/g, "").slice(0, 14);
const parse = (s: string) => Date.parse(
`${s.slice(0,4)}-${s.slice(4,6)}-${s.slice(6,8)}T${s.slice(8,10)}:${s.slice(10,12)}:${s.slice(12,14)}Z`
); The five decisions
Overlap your window
Articles reach the index shortly after they are published, and not all at the same lag. A window that starts exactly where the last one ended will miss whatever arrived late. Overlap by 15 minutes and let deduplication absorb the repeats — the cost is nothing, and the alternative is silent gaps.
Deduplicate on story_id
If you dedup on article id, one wire story becomes forty notifications. Dedup on
story_id and it becomes one, with an outlet count attached that tells the reader how big it is. This is
the single decision that separates a useful monitor from an unusable one.
Bound the map. A story that has been quiet for a day is over; keeping it costs memory and, if it reappears with new coverage, you would rather be told again than not.
Re-notify when a story grows
The first poll that sees a story sees it small: syndication has not propagated, and a wire story that will be on forty domains by lunchtime has three at breakfast. A monitor that notifies once, at first sight, reports every story at its least informative moment and never corrects itself.
So keep the outlet count you reported, and report again when the story crosses a threshold —
5, 15, 40 outlets — or has grown by ten since you last mentioned it. The
/stories row already carries the current outlets, so the comparison is free; the
re-notification itself is one 2-credit call to /stories/{id}, which
returns the current picture — every outlet, first_seen and last_seen, what it is about —
and limit=1 keeps the body small. The steps are geometric on purpose: a story that goes from 40 to 50
outlets is the same story; one that goes from 3 to 15 is news.
Advance only on a complete poll
A failed poll that still advances the watermark loses everything published in that interval, permanently and
silently. So does a successful poll that reached only part of the index: it is 200, it has
rows, and it is missing whatever the unread part held. Advance only when partial is
false and coverage is complete; otherwise keep the watermark where it is and
let the next tick, with its overlap, try again. Reading completeness covers
the values.
Count the stale polls. One incomplete poll is weather; four in a row at a fifteen-minute interval is an hour of not advancing, and someone should know — the loop above alerts at four rather than silently retrying forever.
Pick an interval you can afford
| Interval | Credits / month | Fits |
|---|---|---|
| 1 minute | ~43,200 | Build and up |
| 5 minutes | ~8,640 | Build and up |
| 15 minutes | ~2,880 | Build and up |
| 1 hour | ~720 | Free (1,000/month), with room to spare |
Per query, for a plain poll that costs 1 credit a run; a watch with an entity_id or a
threshold costs more, and every webhook delivery attempt adds 1. Ten monitored topics at 15 minutes is ~28,800
credits a month, more than half of a Build plan's 50,000 before any user has searched for anything. Count them up
front.
A watch at the same interval costs the same requests, charged server-side to the key that created it, without your client making the empty polls: with a webhook, your code runs only when there is something to say. A watch cannot run more often than its plan's floor: 60 minutes on Free, 15 minutes on Build, 5 minutes on Scale, 1 minute on Archive. Plan comparison has the caps.
What to notify on
Not everything new is worth telling someone about. The clustered response gives you the material to decide:
function worthNotifying(story) {
// Broad pickup: many independent outlets, not one syndicated wire.
if (story.outlets >= 8) return true;
// Or narrow but from a publication you care about.
if (story.sources.some((s) => WATCHED.has(s))) return true;
return false;
} outlets is doing the work: it is the difference between "a story broke" and "a press release was
republished". Evaluating coverage goes further.
Backfilling a new monitor
Starting a monitor with a from of a month ago will return a month of stories in one burst. Walk the
month instead, so each request is bounded and the first notification is not a flood.
A monitor defined by filters alone — an organization, a topic, a publisher country — can walk it newest-first with
a cursor. Pages stay disjoint while new articles arrive, so the walk
neither skips nor repeats one, and collapsing on story_id at the end turns articles back into events:
from=$(date -u -d "30 days ago" +%Y-%m-%d)
cursor=start
while [ -n "$cursor" ]; do
page=$(curl -s -G "https://api.unzoi.com/top-headlines" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "organization=maersk" \
--data-urlencode "from=$from" \
--data-urlencode "view=compact" --data-urlencode "limit=100" \
--data-urlencode "cursor=$cursor")
# Stop on an incomplete page; the cursor printed is where to resume.
echo "$page" | jq -e '(.partial | not) and .coverage == "complete"' > /dev/null \
|| { echo "incomplete page; resume with cursor=$cursor" >&2; break; }
echo "$page" | jq -c '.results[]'
cursor=$(echo "$page" | jq -r '.next_cursor // empty')
sleep 0.5
done | jq -s -c 'group_by(.story_id)[] | {story_id: .[0].story_id, articles: length, title: .[0].title}'
A monitor with a q cannot use a cursor. A ranked query has no fixed position to resume from, and
cursor with q is refused. Walk it in day windows instead:
for day in $(seq 30 -1 1); do
from=$(date -u -d "$day days ago" +%Y%m%d000000)
to=$(date -u -d "$((day-1)) days ago" +%Y%m%d000000)
curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
--data-urlencode "q=port congestion" \
--data-urlencode "from=$from" --data-urlencode "to=$to" \
--data-urlencode "limit=100" | jq -c '.stories[]'
sleep 0.5
done
Check your archive depth first — on the free tier either loop silently stops
finding anything past 30 days back, and from_clamped in each response is what tells you so. And read
coverage on each page or window: a day the index answers with range_too_wide has to be
split into smaller windows, because retrying the same request cannot succeed, while timed_out wants a
retry.