← Argus · Methodology · Sources & licensing · Glossary

Track record

A risk score earns trust by being checked against what happened. Argus is young enough that the honest version of this page is mostly method and empty tables: what we will measure, how the record accrues, and the corrections we have already published. It is dated so a reader can see the record fill — or fail to.

What we will measure

Likelihood is an activity index, not a forecast probability (see the methodology). It still makes a checkable claim: a theater in a higher band should, more often than one in a lower band, be followed by the kinds of disruption the register exists to warn about. The calibration check we intend to publish:

Outcomes log

Empty. Daily posture snapshots have accrued since 2026-08-13; the first calibration report is due once a full 120-day window of snapshots exists with its bulk-source backfill complete — December 2026 at the earliest — and will be published here with the query that produced it.

PeriodBandObservationsOutcome rate (14 d)Outcome rate (30 d)
No report yet. Snapshots accruing.

Corrections record

Events withdrawn from the feed

Every event withdrawn from the public feed, with the reason. Withdrawals are soft — the row keeps its id and a retraction_reason, so a subscriber who already pulled it can reconcile against the correction rather than a gap. Reads exclude withdrawn rows everywhere.

This record covers the shared feed only. An organization’s own analyst may overrule an event’s grade for that organization (labelled analyst-adjusted on every surface it touches, with the reason); such overrides never alter the shared row, never appear here, and reach us only as field reports. When one leads us to change the shared feed, that change is recorded above like any other.

DateSourceRowsKindWhat and why
2026-08-28ofac-sdn209privacyOFAC designations that named natural persons were withdrawn from the public feed and the adapter changed to publish pseudonymous individual entries only. A 'removed from the SDN list' delta would otherwise have permanently recorded that a named person used to be sanctioned.
2026-09-05ofac-sdn193data-qualityPhantom re-emissions: a cache-save bug in the daily job (fixed 2026-08-28) re-published the same 85 designations as newly 'added' on five consecutive days. Every repeat after each designation's first appearance was withdrawn with a stated reason; the first appearances stand.

Claims we published that were wrong

The same standard, applied to our own prose. Nothing was retracted from the feed for these — what was wrong was something this site said about the instrument, which is the kind of error a reader has no way to catch. Each entry names the claim as it stood and what replaced it.

DatePageWhat we saidWhat is true
2026-09-16/feedThe feed graded the Houthi seizure of islands in the Bab al-Mandab strait (2026-09-10/11) — front-page, carried by dozens of outlets — as low confidence, the same grade as a one-outlet rumour.The rule was applied as written and the outcome was still wrong: the GDELT export path collapsed every outlet that carried the story into one row and discarded the count, so the promotion rule had no breadth to read. Methodology 2026.09.1 keeps the outlet count (corroboration) and the coder's per-article confidence (coherence) on every machine-coded row, adds the breadth path to medium, and gates rows the coder was unsure of. Replayed from the archive, the seizure rows grade medium. Filed by the analyst as field report #111; the replay is a committed test fixture.
2026-09-07/track-recordThe calibration check defined its outcome as a high-confidence, primary-source event of a disruptive type (armed clash, attack on shipping, missile activity, blockade-class maritime incursion, cable/pipeline incident).No source in the pipeline grades a disruptive-type event high: those five types have never produced a single high-confidence row. The check as published could not have produced a result at any future date. The outcome is redefined above at two tiers (medium-or-better, and machine-coded reported separately), with the resulting weaknesses stated beside it.
2026-09-07/methodologyThe page described how natural hazards are MATCHED (proximity, radius) and never how they are SCORED, while the thresholds it printed were introduced as per-theater figures.A careful reader was therefore led away from the fact that the heaviest hazard weight is below the first likelihood threshold, so a single natural hazard of any severity scores likelihood 1 of 5 and the worst possible single-hazard row is 'medium'. The arithmetic is now stated in full under Likelihood and repeated in Limitations.
2026-09-07/methodologyThe confidence table defined high as 'primary source (official publication, instrument data)', medium as 'reputable secondary source', for every source alike.True of the reporting sources, false of the four hazard adapters, which grade confidence by magnitude, sustained wind, alert colour and alert severity — intensity, not evidential quality. The table now prints both readings side by side and names the consequence: on a hazard row, severity is counted twice in the weight product.
2026-09-07/methodology · /sources“UCDP’s georeferenced conflict data is roughly nine tenths of published events”, and the source register’s coverage note calls UCDP the largest contributor by volume.True of a single pipeline run and roughly true of the whole archive, but not of the 120-day window the scores are computed over: since the GDELT ingest moved to the raw event exports on 2026-08-23, machine-coded records contribute more rows to that window than UCDP does. Both pages now say so; the register string itself is corrected at its next revision.
2026-09-07/methodology“GDELT-sourced candidates are capped at low confidence … and never re-labelled as confirmed.”Never re-labelled as confirmed is true; capped at low confidence is not. The pipeline promotes a low-confidence candidate to medium when a high-confidence event from a different source shares its theater, type and date, which doubles that row's confidence weight. The rule is now described on the methodology page, together with the fact that it partly double-counts a single real-world occurrence, and the calibration check above excludes promoted rows from its curated tier.
2026-09-07/methodologyThe saturation limitation said the theater-likelihood alert kind is disabled because a saturated theater cannot move.It is not disabled. The kind exists and accepts rules; the console states beside such a rule that it cannot fire while its theater is saturated. The limitation now describes what the product actually does.

Methodology versions

Current: argus-methodology-2026.10. A comparison across versions compares two rulers; the product says so wherever it happens.

VersionDateStatusWhat changed
argus-methodology-2026.052026-05shippedInitial scale: confidence × severity weights summed over a 120-day window; fixed likelihood thresholds tuned at fixture scale.
argus-methodology-2026.072026-07shippedDIME instrument dimension; natural-hazard family and proximity matching; disinformation-campaign type. Same likelihood arithmetic.
argus-methodology-2026.092026-09-14shippedLikelihood recalibration, first half: recency-decayed activity (14-day half-life) on geometric bands replaces the undecayed sum on fixed thresholds — because on the live feed every theater had saturated the old scale. Sanctions designations (decision-class) are counted but no longer scored. Adds intensity (the absolute band) and a shaped-but-empty tempo slot. Bands are PROVISIONAL pending analyst review; tempo against each theater's own baseline waits on enough collection history (~Dec 2026).
argus-methodology-2026.09.12026-09-16shippedCorroboration breadth + coherence (E-CAL items 2 and 3, Stage C1). Every event carries `corroboration` (distinct outlets or source adapters that carried it) and machine-coded rows carry `coherence` (GDELT's own best per-article extraction confidence). The promotion rule gains a breadth path: a machine-coded row carried by ≥ 3 distinct outlets grades medium — never high. Rows whose best article scores below 30 are dropped and counted in the health readout; the linked article is the highest-confidence one. Field evidence: the Bab al-Mandab island seizure of 2026-09-10/11 graded low (issue #111); replayed from the archive it now grades medium. Likelihood arithmetic unchanged; promoted rows weigh 0.6 instead of 0.3, so GDELT-heavy theaters read higher than under 2026.09.
argus-methodology-2026.09.22026-09-30shippedCoverage release 1 and a new natural-hazard rule. Shapes: the Middle East widened to all of Iran, the Gulf, Hormuz and the Gulf of Oman; the East China Sea widened north over the Yellow Sea, the Bohai and Beijing–Tianjin; Latin America gets a shape (Mexico through South America and the Caribbean). Country fallback adds the Gulf states and Latin America. Hazards: a site is scored by the strongest hazard ACTIVE on it now (storms and severe weather 3 days, floods 5, earthquakes and wildfires 7, volcanoes 14), counted once per storm or quake and banded by its own weight (≥ 0.8 → 5, ≥ 0.5 → 4, ≥ 0.3 → 3, else 2); a serious hazard of a different kind (a flood beside a cyclone) adds one level — reissued warnings of the same kind never compound; a site keeps a level-2 ‘after the event’ row for 7 days. Replaces an undecayed sum over every in-radius warning of any age, under which any single hazard scored 1 of 5 and piles of old minor warnings saturated. Windows and cut-offs PROVISIONAL pending analyst review. Theater likelihood arithmetic unchanged.
argus-methodology-2026.09.32026-09-30shippedCoverage release 2: Southeast Asia (Myanmar, Thailand, Laos, Cambodia, Malaysia, Singapore, Indonesia, the Philippines, Timor) and South Asia (Afghanistan, Pakistan, India, Nepal, Bhutan, Bangladesh, Sri Lanka, the Maldives, the Arabian Sea) get shapes, drawn along their neighbours' edges so no shape overlaps another. Country fallback adds both regions and Vietnam (South China Sea). UCDP's country routing files Myanmar, Pakistan, Afghanistan, India, Bangladesh, the Philippines, Indonesia and Thailand under the new regions, and the Gulf and Latin American countries the 2026.09.2 shapes reached. Region likelihood arithmetic unchanged.
argus-methodology-2026.102026-10-03currentTempo (Option B). A theater's level is its last 7 days of weighted activity against its own previous 8 weeks (z, banded at -1 / 0.5 / 1.5 / 2.5), so 2 is normal for that theater and an always-busy theater no longer pins the scale. Activity comes from a volume series that measures each theater's GDELT activity over complete UTC days before the feed's per-theater cap (rebuilt from GDELT's public archive), plus the theater's other same-day sources; lagged bulk data and sanctions decisions never count. The previous level stays as intensity, shown beside it. The GDELT cap no longer cuts corroborated rows. Provisional; checked on six weeks of archive data, in which the only reading above 3 was the Red Sea during the Bab al-Mandab week.