← Argus · Methodology · Sources & licensing · Glossary
Track record
A risk score earns trust by being checked against what happened. Argus is young enough that the honest version of this page is mostly method and empty tables: what we will measure, how the record accrues, and the corrections we have already published. It is dated so a reader can see the record fill — or fail to.
What we will measure
Likelihood is an activity index, not a forecast probability (see the methodology). It still makes a checkable claim: a theater in a higher band should, more often than one in a lower band, be followed by the kinds of disruption the register exists to warn about. The calibration check we intend to publish:
- Unit: one (theater, day, band) observation, taken from the daily posture snapshot the feed cron already writes — the same number a customer saw that morning, not a recomputation.
- Outcome (revised 2026-09-07): whether, within the following 14 and 30 days, the theater recorded an event of a disruptive type — armed clash · attack on shipping · missile activity · blockade-class maritime incursion · cable or pipeline incident — drawn from the feed itself, so the check is reproducible from published data. It is counted at two tiers, reported separately and never pooled:
- Tier A — curated or primary. The same types at medium confidence or better. For these types medium is, almost always, the ceiling: the one source that grades a disruptive event high is the NGA maritime warnings, which reach the feed only where a theater covers the water they warn about. The first arrived on 2026-10-01, when the Middle East shape grew to cover the Strait of Hormuz (an in-force mine warning). So Tier A means, in practice, a UCDP conflict record, and now and then an official maritime warning.
- Tier B — machine-coded tripwire. The same types at any confidence, which in practice means GDELT’s low-confidence records. Reported beside Tier A and never merged into it, because in most theaters those same records are also driving the score being checked — Tier B is an internal-consistency check, not a validation.
- Report: per band and per tier, the observed outcome rate with a confidence interval, per theater and pooled; the same for the intensity band once recalibration ships. Bands that do not separate are a finding, not a footnote.
- Known weaknesses of the check, stated up front:
- It is not independent of the thing it checks. The outcome is drawn from the same feed that drives the score, so a source going quiet lowers both.
- Outcomes arrive late. The bulk conflict source lags 42–120 days, so outcome rates for recent windows stay provisional until its backfill lands.
- A permanently violent theater passes without predicting anything. It will score “high band, high outcome rate” on base rates alone — which is exactly why the recalibration measures tempo against the theater’s own baseline.
- Tier A is unobservable in five of the twelve geographic theaters. In the window this site ships today, the Taiwan Strait, South China Sea, East China Sea, Korean Peninsula and Baltic & Nordic record no disruptive-type event from a curated or primary source: every disruptive row in them is machine-coded, and the only ones above low confidence are breadth promotions, which Tier A excludes (next item). Those are the coercion flashpoints the product is bought for, so the first report will be able to check the active-conflict theaters and not those. Closing the gap needs a primary-source disruptive feed. The maritime-warning adapter (NGA MSI, listed on the sources page) shipped on 2026-09-16 and would grade such a row high — but 51 of the 54 security warnings in force fall outside every theater polygon, so it has produced no disruptive-type row in these theaters yet. The theaters are drawn too narrowly for the warnings that matter. Methodology 2026.09.2 widened the Middle East over the Gulf and Hormuz; the rest of the geometry proposal is still to be decided. Southeast Asia and South Asia gained shapes on 2026-09-30 (2026.09.3); their curated conflict records make both observable in Tier A from the first publish, so their record starts on 2026-10-01.
- Confidence is not purely a source property, so the two tiers leak. The pipeline promotes a machine-coded candidate from low to medium when a high-confidence event from another source shares its theater, type and date, and — since methodology 2026.09.1 — when enough distinct outlets carried it. Such a row would satisfy Tier A while being a Tier B row wearing a promotion, so the check must exclude events tagged
corroborated-by:from Tier A — otherwise it counts its own tripwires as outcomes. The breadth path is not rare (it lifts roughly a third of a day’s machine-coded rows), but every promotion is tagged, which is what makes excluding it possible; the exclusion is now what keeps the five theaters above honestly listed as blind. - Even Tier A means “a dataset recorded it”, not “a government confirmed it”. Its rows are curated conflict records coded from reports weeks after the fact. The stronger evidential standard the first definition reached for is the right one; the feed cannot currently supply it, and saying so is more useful than publishing a check that silently never runs.
Outcomes log
Empty. Daily posture snapshots have accrued since 2026-08-13; the first calibration report is due once a full 120-day window of snapshots exists with its bulk-source backfill complete — December 2026 at the earliest — and will be published here with the query that produced it.
| Period | Band | Observations | Outcome rate (14 d) | Outcome rate (30 d) |
|---|---|---|---|---|
| No report yet. Snapshots accruing. | ||||
Corrections record
Events withdrawn from the feed
Every event withdrawn from the public feed, with the reason. Withdrawals are soft — the row keeps its id and a retraction_reason, so a subscriber who already pulled it can reconcile against the correction rather than a gap. Reads exclude withdrawn rows everywhere.
This record covers the shared feed only. An organization’s own analyst may overrule an event’s grade for that organization (labelled analyst-adjusted on every surface it touches, with the reason); such overrides never alter the shared row, never appear here, and reach us only as field reports. When one leads us to change the shared feed, that change is recorded above like any other.
| Date | Source | Rows | Kind | What and why |
|---|---|---|---|---|
| 2026-08-28 | ofac-sdn | 209 | privacy | OFAC designations that named natural persons were withdrawn from the public feed and the adapter changed to publish pseudonymous individual entries only. A 'removed from the SDN list' delta would otherwise have permanently recorded that a named person used to be sanctioned. |
| 2026-09-05 | ofac-sdn | 193 | data-quality | Phantom re-emissions: a cache-save bug in the daily job (fixed 2026-08-28) re-published the same 85 designations as newly 'added' on five consecutive days. Every repeat after each designation's first appearance was withdrawn with a stated reason; the first appearances stand. |
Claims we published that were wrong
The same standard, applied to our own prose. Nothing was retracted from the feed for these — what was wrong was something this site said about the instrument, which is the kind of error a reader has no way to catch. Each entry names the claim as it stood and what replaced it.
| Date | Page | What we said | What is true |
|---|---|---|---|
| 2026-09-16 | /feed | The feed graded the Houthi seizure of islands in the Bab al-Mandab strait (2026-09-10/11) — front-page, carried by dozens of outlets — as low confidence, the same grade as a one-outlet rumour. | The rule was applied as written and the outcome was still wrong: the GDELT export path collapsed every outlet that carried the story into one row and discarded the count, so the promotion rule had no breadth to read. Methodology 2026.09.1 keeps the outlet count (corroboration) and the coder's per-article confidence (coherence) on every machine-coded row, adds the breadth path to medium, and gates rows the coder was unsure of. Replayed from the archive, the seizure rows grade medium. Filed by the analyst as field report #111; the replay is a committed test fixture. |
| 2026-09-07 | /track-record | The calibration check defined its outcome as a high-confidence, primary-source event of a disruptive type (armed clash, attack on shipping, missile activity, blockade-class maritime incursion, cable/pipeline incident). | No source in the pipeline grades a disruptive-type event high: those five types have never produced a single high-confidence row. The check as published could not have produced a result at any future date. The outcome is redefined above at two tiers (medium-or-better, and machine-coded reported separately), with the resulting weaknesses stated beside it. |
| 2026-09-07 | /methodology | The page described how natural hazards are MATCHED (proximity, radius) and never how they are SCORED, while the thresholds it printed were introduced as per-theater figures. | A careful reader was therefore led away from the fact that the heaviest hazard weight is below the first likelihood threshold, so a single natural hazard of any severity scores likelihood 1 of 5 and the worst possible single-hazard row is 'medium'. The arithmetic is now stated in full under Likelihood and repeated in Limitations. |
| 2026-09-07 | /methodology | The confidence table defined high as 'primary source (official publication, instrument data)', medium as 'reputable secondary source', for every source alike. | True of the reporting sources, false of the four hazard adapters, which grade confidence by magnitude, sustained wind, alert colour and alert severity — intensity, not evidential quality. The table now prints both readings side by side and names the consequence: on a hazard row, severity is counted twice in the weight product. |
| 2026-09-07 | /methodology · /sources | “UCDP’s georeferenced conflict data is roughly nine tenths of published events”, and the source register’s coverage note calls UCDP the largest contributor by volume. | True of a single pipeline run and roughly true of the whole archive, but not of the 120-day window the scores are computed over: since the GDELT ingest moved to the raw event exports on 2026-08-23, machine-coded records contribute more rows to that window than UCDP does. Both pages now say so; the register string itself is corrected at its next revision. |
| 2026-09-07 | /methodology | “GDELT-sourced candidates are capped at low confidence … and never re-labelled as confirmed.” | Never re-labelled as confirmed is true; capped at low confidence is not. The pipeline promotes a low-confidence candidate to medium when a high-confidence event from a different source shares its theater, type and date, which doubles that row's confidence weight. The rule is now described on the methodology page, together with the fact that it partly double-counts a single real-world occurrence, and the calibration check above excludes promoted rows from its curated tier. |
| 2026-09-07 | /methodology | The saturation limitation said the theater-likelihood alert kind is disabled because a saturated theater cannot move. | It is not disabled. The kind exists and accepts rules; the console states beside such a rule that it cannot fire while its theater is saturated. The limitation now describes what the product actually does. |
Methodology versions
Current: argus-methodology-2026.10. A comparison across versions compares two rulers; the product says so wherever it happens.
| Version | Date | Status | What changed |
|---|---|---|---|
argus-methodology-2026.05 | 2026-05 | shipped | Initial scale: confidence × severity weights summed over a 120-day window; fixed likelihood thresholds tuned at fixture scale. |
argus-methodology-2026.07 | 2026-07 | shipped | DIME instrument dimension; natural-hazard family and proximity matching; disinformation-campaign type. Same likelihood arithmetic. |
argus-methodology-2026.09 | 2026-09-14 | shipped | Likelihood recalibration, first half: recency-decayed activity (14-day half-life) on geometric bands replaces the undecayed sum on fixed thresholds — because on the live feed every theater had saturated the old scale. Sanctions designations (decision-class) are counted but no longer scored. Adds intensity (the absolute band) and a shaped-but-empty tempo slot. Bands are PROVISIONAL pending analyst review; tempo against each theater's own baseline waits on enough collection history (~Dec 2026). |
argus-methodology-2026.09.1 | 2026-09-16 | shipped | Corroboration breadth + coherence (E-CAL items 2 and 3, Stage C1). Every event carries `corroboration` (distinct outlets or source adapters that carried it) and machine-coded rows carry `coherence` (GDELT's own best per-article extraction confidence). The promotion rule gains a breadth path: a machine-coded row carried by ≥ 3 distinct outlets grades medium — never high. Rows whose best article scores below 30 are dropped and counted in the health readout; the linked article is the highest-confidence one. Field evidence: the Bab al-Mandab island seizure of 2026-09-10/11 graded low (issue #111); replayed from the archive it now grades medium. Likelihood arithmetic unchanged; promoted rows weigh 0.6 instead of 0.3, so GDELT-heavy theaters read higher than under 2026.09. |
argus-methodology-2026.09.2 | 2026-09-30 | shipped | Coverage release 1 and a new natural-hazard rule. Shapes: the Middle East widened to all of Iran, the Gulf, Hormuz and the Gulf of Oman; the East China Sea widened north over the Yellow Sea, the Bohai and Beijing–Tianjin; Latin America gets a shape (Mexico through South America and the Caribbean). Country fallback adds the Gulf states and Latin America. Hazards: a site is scored by the strongest hazard ACTIVE on it now (storms and severe weather 3 days, floods 5, earthquakes and wildfires 7, volcanoes 14), counted once per storm or quake and banded by its own weight (≥ 0.8 → 5, ≥ 0.5 → 4, ≥ 0.3 → 3, else 2); a serious hazard of a different kind (a flood beside a cyclone) adds one level — reissued warnings of the same kind never compound; a site keeps a level-2 ‘after the event’ row for 7 days. Replaces an undecayed sum over every in-radius warning of any age, under which any single hazard scored 1 of 5 and piles of old minor warnings saturated. Windows and cut-offs PROVISIONAL pending analyst review. Theater likelihood arithmetic unchanged. |
argus-methodology-2026.09.3 | 2026-09-30 | shipped | Coverage release 2: Southeast Asia (Myanmar, Thailand, Laos, Cambodia, Malaysia, Singapore, Indonesia, the Philippines, Timor) and South Asia (Afghanistan, Pakistan, India, Nepal, Bhutan, Bangladesh, Sri Lanka, the Maldives, the Arabian Sea) get shapes, drawn along their neighbours' edges so no shape overlaps another. Country fallback adds both regions and Vietnam (South China Sea). UCDP's country routing files Myanmar, Pakistan, Afghanistan, India, Bangladesh, the Philippines, Indonesia and Thailand under the new regions, and the Gulf and Latin American countries the 2026.09.2 shapes reached. Region likelihood arithmetic unchanged. |
argus-methodology-2026.10 | 2026-10-03 | current | Tempo (Option B). A theater's level is its last 7 days of weighted activity against its own previous 8 weeks (z, banded at -1 / 0.5 / 1.5 / 2.5), so 2 is normal for that theater and an always-busy theater no longer pins the scale. Activity comes from a volume series that measures each theater's GDELT activity over complete UTC days before the feed's per-theater cap (rebuilt from GDELT's public archive), plus the theater's other same-day sources; lagged bulk data and sanctions decisions never count. The previous level stays as intensity, shown beside it. The GDELT cap no longer cuts corroborated rows. Provisional; checked on six weeks of archive data, in which the only reading above 3 was the Red Sea during the Bab al-Mandab week. |