Skip to content
unzoi docs
Search and navigation
Start here
REST API
MCP
Limits and plans
Agent clients
SDKs
Guides

Monitoring a topic

You can poll, or you can let the index poll for you. Either way, five things make a monitor correct: a watermark, a deduplication key, a rule for when a story you already reported is worth mentioning again, a check that the poll was complete before the watermark moves, and an interval you can afford.

Create a watch instead

A watch is the loop on this page, run by the index against your key. It runs a /stories query every interval_minutes, overlaps each window with the last, remembers stories by story_id, reports a story again when it crosses 5, 15, 40 or 100 outlets or grows by ten, and moves its watermark only on a complete run. Four of the five decisions are made for you. The interval is still yours, and your plan sets its floor.

curl -s -X POST "https://api.unzoi.com/watches" -H "x-api-key: $UNZOI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "label": "port congestion",
    "q": "port congestion OR container backlog",
    "interval_minutes": 15,
    "minimum_outlets": 3,
    "webhook_url": "https://hooks.example.com/unzoi"
  }'

The response carries the watch's id and, because it has a webhook_url, its webhook_secret. The secret is returned once. Store it then. From an agent, the same request is the create_watch tool.

Reading what it found

Every event lands in the watch's feed, webhook or not: new_story when a story first reaches minimum_outlets, story_growth when it grows. Read it with GET /watches/{id}/events, or the watch_events tool.

curl -s "https://api.unzoi.com/watches/$WATCH_ID/events?limit=20" -H "x-api-key: $UNZOI_KEY" \
  | jq -r '.events[] | "\(.kind)\t\(.story.outlets)\t\(.story.title)"'

Events are newest first and kept seven days. The feed is in the order events were recorded: each id begins with when the run that recorded it was due. Read from the top until you reach an id you have already handled, and stop there. Each page you read is a request. label names the watch in your lists; name in a watch is the proper-name filter, as on /stories.

Verifying a webhook

With a webhook_url, each event is also POSTed to you. Check x-unzoi-signature against the raw body before trusting it, answer 2xx within ten seconds, and deduplicate on x-unzoi-delivery: a failed delivery is sent again after 1, 5 and 25 minutes.

import { createHmac, timingSafeEqual } from "node:crypto";

// rawBody is the request body exactly as received, before JSON.parse.
function verify(rawBody: Buffer, header: string | undefined, secret: string): boolean {
  const expected = Buffer.from("sha256=" + createHmac("sha256", secret).update(rawBody).digest("hex"));
  const received = Buffer.from(header ?? "");
  return received.length === expected.length && timingSafeEqual(received, expected);
}

Watches has the Python version, the payload, and what happens when deliveries keep failing.

When to poll instead

The loop below is the portable path. Keep it when:

  • you run the index yourself. Watches are served by the hosted API; the local stdio server answers the watch tools with unsupported.
  • the query needs something a watch does not carry: its own time window, a sort, signal_percentile_min.
  • the growth rule should be yours rather than the watch's steps.
  • one interval matches more than 100 stories. A watch's run reads the newest 100 in its window.

The loop

import { setTimeout as sleep } from "node:timers/promises";

const BASE = "https://api.unzoi.com";
const headers = { "x-api-key": process.env.UNZOI_KEY! };

// The watermark. Persist it — an in-memory one restarts your feed from scratch
// on every deploy, which is how a monitor floods a channel at 3am.
let watermark = loadWatermark() ?? stamp(Date.now() - 3600_000);
// What you have already told someone about, and how big it was then.
const known = new Map<string, { outlets: number; count: number; notifiedAt: number }>(loadKnown());
let stalePolls = 0;

async function poll() {
  const url = new URL(BASE + "/stories");
  url.searchParams.set("q", "port congestion OR container backlog");
  url.searchParams.set("mode", "hybrid");
  // Overlap deliberately: articles are indexed a little after they are
  // published, so a watermark exactly at the last poll misses the stragglers.
  url.searchParams.set("from", stamp(parse(watermark) - 15 * 60_000));
  url.searchParams.set("limit", "100");

  const response = await fetch(url, { headers });
  if (response.status === 429) {
    if ((await response.clone().json()).code === "insufficient_credits") {
      console.error("out of credits — pausing until the month resets");
      return;
    }
    await sleep(1000);
    return;
  }
  if (!response.ok) return;   // transient; the next tick catches up

  const page = await response.json();

  for (const story of page.stories) {
    const before = known.get(story.story_id);
    if (!before) {
      // New event. Dedup on story_id, not on article id: forty outlets running
      // one wire story is ONE thing that happened, and forty notifications is spam.
      known.set(story.story_id, { outlets: story.outlets, count: story.count, notifiedAt: Date.now() });
      notify(story, "new");
    } else if (grew(before.outlets, story.outlets)) {
      // Something you already reported got bigger. Expand it for the current
      // picture — 2 credits, and limit=1 keeps the body small.
      const detail = await (await fetch(`${BASE}/stories/${story.story_id}?limit=1`, { headers })).json();
      known.set(story.story_id, { outlets: detail.outlets, count: detail.count, notifiedAt: Date.now() });
      notify(detail, "growth");
    }
  }

  // Advance only on a complete poll. A poll that reached part of the index is
  // not a poll of that interval, and advancing on it opens a silent gap.
  if (page.partial || page.coverage !== "complete") {
    if (++stalePolls >= 4) alertOperator(`${stalePolls} incomplete polls in a row (${page.coverage})`);
    return;
  }
  stalePolls = 0;
  watermark = stamp(Date.now());
  saveWatermark(watermark);
  saveKnown(prune(known));
}

// A story is worth mentioning again when its outlet count crosses a step, or
// has grown by ten since you last mentioned it.
const STEPS = [5, 15, 40];
const grew = (before: number, now: number) =>
  now - before >= 10 || STEPS.some((step) => before < step && now >= step);

// Bound the map: a story quiet for a day is over, and an unbounded map is a
// memory leak with a slow fuse.
const prune = (m: typeof known) =>
  new Map([...m].filter(([, s]) => Date.now() - s.notifiedAt < 24 * 3600_000));

for (;;) {
  await poll();
  await sleep(15 * 60_000);
}

const stamp = (ms: number) => new Date(ms).toISOString().replace(/[-:T]/g, "").slice(0, 14);
const parse = (s: string) => Date.parse(
  `${s.slice(0,4)}-${s.slice(4,6)}-${s.slice(6,8)}T${s.slice(8,10)}:${s.slice(10,12)}:${s.slice(12,14)}Z`
);

The five decisions

Overlap your window

Articles reach the index shortly after they are published, and not all at the same lag. A window that starts exactly where the last one ended will miss whatever arrived late. Overlap by 15 minutes and let deduplication absorb the repeats — the cost is nothing, and the alternative is silent gaps.

Deduplicate on story_id

If you dedup on article id, one wire story becomes forty notifications. Dedup on story_id and it becomes one, with an outlet count attached that tells the reader how big it is. This is the single decision that separates a useful monitor from an unusable one.

Bound the map. A story that has been quiet for a day is over; keeping it costs memory and, if it reappears with new coverage, you would rather be told again than not.

Re-notify when a story grows

The first poll that sees a story sees it small: syndication has not propagated, and a wire story that will be on forty domains by lunchtime has three at breakfast. A monitor that notifies once, at first sight, reports every story at its least informative moment and never corrects itself.

So keep the outlet count you reported, and report again when the story crosses a threshold — 5, 15, 40 outlets — or has grown by ten since you last mentioned it. The /stories row already carries the current outlets, so the comparison is free; the re-notification itself is one 2-credit call to /stories/{id}, which returns the current picture — every outlet, first_seen and last_seen, what it is about — and limit=1 keeps the body small. The steps are geometric on purpose: a story that goes from 40 to 50 outlets is the same story; one that goes from 3 to 15 is news.

Advance only on a complete poll

A failed poll that still advances the watermark loses everything published in that interval, permanently and silently. So does a successful poll that reached only part of the index: it is 200, it has rows, and it is missing whatever the unread part held. Advance only when partial is false and coverage is complete; otherwise keep the watermark where it is and let the next tick, with its overlap, try again. Reading completeness covers the values.

Count the stale polls. One incomplete poll is weather; four in a row at a fifteen-minute interval is an hour of not advancing, and someone should know — the loop above alerts at four rather than silently retrying forever.

Pick an interval you can afford

IntervalCredits / monthFits
1 minute~43,200Build and up
5 minutes~8,640Build and up
15 minutes~2,880Build and up
1 hour~720Free (1,000/month), with room to spare

Per query, for a plain poll that costs 1 credit a run; a watch with an entity_id or a threshold costs more, and every webhook delivery attempt adds 1. Ten monitored topics at 15 minutes is ~28,800 credits a month, more than half of a Build plan's 50,000 before any user has searched for anything. Count them up front.

A watch at the same interval costs the same requests, charged server-side to the key that created it, without your client making the empty polls: with a webhook, your code runs only when there is something to say. A watch cannot run more often than its plan's floor: 60 minutes on Free, 15 minutes on Build, 5 minutes on Scale, 1 minute on Archive. Plan comparison has the caps.

What to notify on

Not everything new is worth telling someone about. The clustered response gives you the material to decide:

function worthNotifying(story) {
  // Broad pickup: many independent outlets, not one syndicated wire.
  if (story.outlets >= 8) return true;
  // Or narrow but from a publication you care about.
  if (story.sources.some((s) => WATCHED.has(s))) return true;
  return false;
}

outlets is doing the work: it is the difference between "a story broke" and "a press release was republished". Evaluating coverage goes further.

Backfilling a new monitor

Starting a monitor with a from of a month ago will return a month of stories in one burst. Walk the month instead, so each request is bounded and the first notification is not a flood.

A monitor defined by filters alone — an organization, a topic, a publisher country — can walk it newest-first with a cursor. Pages stay disjoint while new articles arrive, so the walk neither skips nor repeats one, and collapsing on story_id at the end turns articles back into events:

from=$(date -u -d "30 days ago" +%Y-%m-%d)
cursor=start
while [ -n "$cursor" ]; do
  page=$(curl -s -G "https://api.unzoi.com/top-headlines" -H "x-api-key: $UNZOI_KEY" \
    --data-urlencode "organization=maersk" \
    --data-urlencode "from=$from" \
    --data-urlencode "view=compact" --data-urlencode "limit=100" \
    --data-urlencode "cursor=$cursor")
  # Stop on an incomplete page; the cursor printed is where to resume.
  echo "$page" | jq -e '(.partial | not) and .coverage == "complete"' > /dev/null \
    || { echo "incomplete page; resume with cursor=$cursor" >&2; break; }
  echo "$page" | jq -c '.results[]'
  cursor=$(echo "$page" | jq -r '.next_cursor // empty')
  sleep 0.5
done | jq -s -c 'group_by(.story_id)[] | {story_id: .[0].story_id, articles: length, title: .[0].title}'

A monitor with a q cannot use a cursor. A ranked query has no fixed position to resume from, and cursor with q is refused. Walk it in day windows instead:

for day in $(seq 30 -1 1); do
  from=$(date -u -d "$day days ago" +%Y%m%d000000)
  to=$(date -u -d "$((day-1)) days ago" +%Y%m%d000000)
  curl -s -G "https://api.unzoi.com/stories" -H "x-api-key: $UNZOI_KEY" \
    --data-urlencode "q=port congestion" \
    --data-urlencode "from=$from" --data-urlencode "to=$to" \
    --data-urlencode "limit=100" | jq -c '.stories[]'
  sleep 0.5
done

Check your archive depth first — on the free tier either loop silently stops finding anything past 30 days back, and from_clamped in each response is what tells you so. And read coverage on each page or window: a day the index answers with range_too_wide has to be split into smaller windows, because retrying the same request cannot succeed, while timed_out wants a retry.