← Argus · Sources & licensing · Glossary · Track record · Live demo
Scoring methodology
Argus produces structured evidence, not analysis: every score below is a deterministic computation over cited events, and every constant on this page is rendered from the module that computes with it — the page cannot drift from the code. The same engine runs twice, in TypeScript and Python, held identical by a cross-language parity suite.
Likelihood (1–5)
Likelihood is tempo (since methodology 2026.10): a theater’s last 7 days of weighted activity against its own previous 8 weeks, z = (now − norm mean) / max(norm spread, 10% of the norm mean), banded at ≥ -1 → 2 · ≥ 0.5 → 3 · ≥ 1.5 → 4 · ≥ 2.5 → 5 (below -1 → 1). So 2 is normal for that theater, 1 is quieter than usual and 4–5 is well above its usual level — a theater that is always busy reads 2, and a genuine surge stands out whatever the theater’s volume.
Daily activity comes from two places. The volume series measures each theater’s machine-coded (GDELT) activity over every complete UTC day before the feed’s per-theater cap, so a busy theater’s surge is not hidden by how many rows the feed keeps; it is rebuilt from GDELT’s public archive with the same filters. To that we add the theater’s other same-day sources (official maritime warnings, Taiwan’s defence ministry, advisories, your own private events, and an exercise’s injected events). Lagged bulk conflict data and sanctions decisions never count toward tempo. A theater with fewer than 28 days of norm, or a series more than 3 days old, keeps the intensity band below as its level, and says so.
Intensity is the second number, shown beside it: how loud the theater is on one scale shared by every theater. Per theater, over a 120-day event window: activity = Σ confidence_weight × severity_weight × 0.5^(age / 14 d) across the theater’s events, then likelihood = 1 + (bands cleared). Age is whole days from the event to the computation date, so an event from this morning counts in full and one 14 days old counts half (recency decay, since methodology 2026.09). Both weight tables and the intensity bands are below, in full. Every score cites its driving events by id — the score is an index into evidence, not a judgment.
provisional The tempo edges, the norm length and the intensity bands are our provisional judgement: what “4 of 5” should mean to a continuity planner is a doctrinal call, labeled as such on every surface that renders the number and open to analyst review. A change moves the constants, never the shape.
Counted, never scored: events from decision-class sources (ofac-sdn, federal-register-bis) — a sanctions designation is a decision, not an incident — appear in a theater’s event count and in the feed, and contribute nothing to its activity.
Confidence weights
One column of this table used to say “meaning”. It had one meaning for the reporting sources and a different one for the four hazard adapters, which is a defect, so both are now printed side by side.
| Confidence | Weight | Reporting sources (OFAC · Taiwan MND · NCSC · IC3 · UCDP · GDELT) | Hazard adapters (USGS · GDACS · NHC · NWS) |
|---|---|---|---|
| high | 1 | primary source — an official publication or instrument record | intensity, not evidence: magnitude ≥ 6 · sustained wind ≥ 64 kt (hurricane strength) · GDACS Red · NWS Extreme |
| medium | 0.6 | reputable secondary source, or a curated dataset coded after the fact | magnitude 4.5–6 · wind 34–64 kt (tropical storm) · GDACS Orange · NWS Severe |
| low | 0.3 | uncorroborated — news-derived candidates enter here. Two rules can promote one to medium, never higher: a primary source reporting the same thing, or enough distinct outlets carrying it (see Limitations) | anything weaker than the bands above |
For hazards, “confidence” is intensity, not evidential quality — so severity is counted twice. All four hazard adapters grade confidence by how strong the phenomenon is: USGS by magnitude, NHC by sustained wind, NWS by the alert’s own severity field, GDACS by alert colour. Nothing in that grading is about whether the event happened. An M4.5 earthquake is instrument-measured and published by the USGS — as certain as anything in this feed — and it is graded medium and discounted to 0.6. The consequence is structural: on a hazard row the two factors of confidence × severity are not independent, both are reading intensity, and no hazard row’s weight can be read as “how sure are we that this occurred”. Separating an evidential axis from an intensity axis is an engine change in both engines with a parity regeneration behind it; it is recorded and not yet done. Until it is, read hazard confidence as a second severity column.
Severity weights (by event type)
| Event type | Weight | Instrument (DIME) |
|---|---|---|
| armed-clash | 1 | Military |
| missile-activity | 1 | Military |
| attack-on-shipping | 1 | Military |
| tropical-cyclone | 0.9 | — |
| earthquake | 0.8 | — |
| volcanic | 0.8 | — |
| adiz-incursion | 0.7 | Military |
| maritime-incursion | 0.7 | Military |
| military-exercise | 0.7 | Military |
| flood | 0.7 | — |
| wildfire | 0.7 | — |
| cable-pipeline-incident | 0.6 | Military |
| cyber-operation | 0.6 | Informational |
| naval-transit | 0.5 | Military |
| gps-interference | 0.5 | Informational |
| severe-weather | 0.5 | — |
| labor-action | 0.5 | — |
| disinformation-campaign | 0.4 | Informational |
| sanctions-action | 0.4 | Economic |
| export-control-action | 0.4 | Economic |
| economic-coercion | 0.4 | Economic |
| diplomatic-incident | 0.3 | Diplomatic |
| protest-unrest | 0.3 | Diplomatic |
| other | 0.3 | — |
| political-statement | 0.2 | Diplomatic |
Intensity bands
Decayed activity ≥ 2 → 2 · ≥ 8 → 3 · ≥ 32 → 4 · ≥ 96 → 5 — geometric, each band four times the last, so one more band means four times the recent weighted activity. These are provisional expert-judgment constants, not derivations — see Limitations, which is not a footnote. The hazard path below does not use them.
Natural hazards: the strongest active hazard sets the level
A hazard is a point-and-radius phenomenon, not a theater one, so a radius-bearing event is removed from theater activity entirely and scored per asset, on its own rule (provisional, like the bands above):
- Only hazards active now count. A report stays active for 7 days (
earthquake), 5 days (flood), 3 days (tropical-cyclone), 7 days (wildfire), 3 days (severe-weather) and 14 days (volcanic). When the last one ends, the site keeps an after the event row at level 2 for 7 more days — recovery, not a threat. - One storm counts once. Repeated reports of the same event (same type and title) collapse to the strongest of them.
- The strongest hazard sets the level, from its own weight
confidence × severity: ≥ 0.8 → 5 · ≥ 0.5 → 4 · ≥ 0.3 → 3 · below → 2. - A serious hazard of a different kind adds one level. When another kind of hazard (a flood beside a cyclone, say) is active on the same site at level 3 or above, the level rises by one, capped at 5. More reports of the same kind add nothing — sources reissue one ongoing warning many times — so a pile of minor warnings no longer adds up.
- So a high-confidence
tropical-cyclone(weight 0.9) over a criticality-5 site scores 5 × 5 = 25 — critical.
| Hazard | Severity | Level at high / medium / low confidence | Active for |
|---|---|---|---|
earthquake | 0.8 | 5 / 3 / 2 | 7 days |
flood | 0.7 | 4 / 3 / 2 | 5 days |
tropical-cyclone | 0.9 | 5 / 4 / 2 | 3 days |
wildfire | 0.7 | 4 / 3 / 2 | 7 days |
severe-weather | 0.5 | 4 / 3 / 2 | 3 days |
volcanic | 0.8 | 5 / 3 / 2 | 14 days |
Why it changed (2026.09.2). The previous rule summed every in-radius hazard of any age, undecayed, against absolute thresholds (≥ 1 → 2 · ≥ 3 → 3 · ≥ 6.5 → 4 · ≥ 11 → 5) set for whole regions. It failed in both directions: the heaviest single hazard weighed less than the first threshold, so one hazard of any severity scored 1 of 5 — a live hurricane over a site read 1 — while a site inside many old, minor warnings saturated at 5. On the live snapshot, 76% of hazard rows were more than a week old. The level still reads the hazard’s reported strength and whether the site is inside its footprint, not conditions at the site; see Limitations.
Impact (1–5), risk, and propagation
Impact is the owner-assigned asset criticality (1–5) — a Business Impact Analysis input, not a computed value. A susceptibility tag raises impact by one (capped at 5) while a matching acute hazard is active — the chronic condition (a flood zone) is always true; the storm is now. Tags and the hazards that trigger them: seismic_zone ← earthquake · flood_zone ← flood · coastal_surge ← tropical-cyclone, severe-weather · wildfire_interface ← wildfire.
Risk = likelihood × impact on a 5×5 matrix: low <5 · medium <10 · high <15 · critical ≥15. Supply-chain propagation is one labeled hop: a downstream asset inherits 0.8 of an upstream’s likelihood when that exceeds its own direct exposure, always labeled inherited via <id> — one honest hop, not a multi-hop digital twin. Natural hazards match by proximity instead of theater: great-circle distance (haversine) against the event’s published radius.
Presentation rules — where honesty is enforced
Three rules govern how scores are shown, because a scale’s floor and ceiling are not measurements:
- Floor: a theater with fewer than 3 contributing events renders “insufficient data”, never “1/5” — “we barely looked” must not read as “we looked, risk is low”.
- Ceiling: a theater still scored on the absolute intensity band (no tempo norm yet) whose activity is ≥ 3× past the top band (288+) renders “saturated” — the 5 is a ceiling, not a ranking. A tempo 5 is never saturated: it means far above that theater’s usual level.
- Age: every score surface shows the age of its newest contributing event, with “stale” marked beyond 7 days — a score is only as current as its freshest evidence, and sources publish with lags from minutes to months.
Limitations — stated by us, not discovered by you
- 5×5 matrices have known pathologies. Multiplying ordinal scales compresses range and can rank-reverse (Cox, What’s Wrong with Risk Matrices?, Risk Analysis 28(2), 2008). We use the matrix because it is the lingua franca of the frameworks our customers are audited against (ISO 31000/22301, TARA), and we compensate by citing drivers on every cell rather than asking anyone to trust the cell.
- Tempo is provisional, and it measures a theater against itself. Until 2026-09-14 the scale was an undecayed sum that pegged every theater at 5; recency decay and geometric bands separated them; by 2026-10 five of thirteen were pegged again, because the busiest theaters all run at the feed’s per-theater cap. Tempo (methodology 2026.10) compares each theater with its own norm instead, from a volume series measured before the cap. Two consequences to read with it: a theater that is permanently violent reads 2, “normal for here” — its intensity says how loud it is; and the machine-coded series measures what GDELT codes, so a change in GDELT’s own coverage moves tempo too. On the six weeks it was checked against (2026-08-15 to 10-01), the only theater that read above 3 on any day was the Red Sea during the Bab al-Mandab seizure week; whether it would miss a quieter escalation is not something six weeks can show. The edges and the norm length are our judgement and ship labeled provisional. Each theater’s z is in the API as
tempo_z(null where the theater fell back to intensity), and lagged bulk data — a latency class of its own — never counts toward it. The theater-likelihood alert kind is not disabled; the alert baseline is reset at each version cutover so the re-banding itself does not fire a storm of “changes”. - Hazard levels are provisional and read the report, not the site. The cut-offs (≥ 0.8 → 5 · ≥ 0.5 → 4 · ≥ 0.3 → 3 · below → 2), the one-level step for a second kind of hazard and the active windows are a first calibration and an analyst’s call, like the theater bands. A level comes from the hazard’s reported strength and whether the site sits inside its footprint; it does not model distance to the eye, flood depth or building standards. The rule and the one it replaced are above, under Likelihood.
- Hazard confidence is intensity, not evidential quality. The four hazard adapters grade confidence by magnitude, wind speed, alert colour and alert severity, so on a hazard row severity is counted twice in
confidence × severityand a well-attested moderate hazard is discounted for being moderate. See the confidence table above, where both readings are printed side by side. - Confidence weighting conflates two things — whether an event happened and how much it matters. A low-confidence armed clash and a high-confidence statement can score similarly. The weights are disclosed above precisely so this is inspectable.
- Propagation is one hop at a fixed damping. Deep dependency chains understate inherited risk by design until multi-hop, gate-aware inheritance ships — we prefer an honest single hop over an unvalidated tree.
- Two sources supply most of the volume, and neither is a primary source. UCDP’s curated conflict records and GDELT’s machine-coded news records are the two largest contributors to the 120-day scoring window by a wide margin. UCDP is coded from reports long after the fact — the source register puts its lag at 42–120 days, and in the snapshot this site ships nothing from it is anywhere near the 7-day freshness mark — so it is excellent evidence about what has happened and none at all about this morning. GDELT is minutes old and enters at low confidence (the one rule that can promote it is stated below). Correction (2026-09-07): this bullet previously said UCDP was “roughly nine tenths of published events”. That was true of a single pipeline run, and remains roughly true of the whole archive, but it is not true of the 120-day window the scores are actually computed over — since the GDELT ingest moved to the raw event exports on 2026-08-23, machine-coded records have contributed more rows to that window than UCDP does. Recorded on the track record; see sources & licensing for the per-source cadence and lag.
- Decay is one global half-life, applied to every source alike. A 14-day half-life is the same order as the freshness window and, on the live archive, moves the distribution by at most one band across 3–21 days — so the choice is not load-bearing for this half of the recalibration. It is load-bearing for tempo, where it cannot yet be fitted. And because the decay reads the event date, a source that publishes weeks after the fact (UCDP, above) is discounted for its lag as if the world had gone quiet — the honest reading is “we are not watching this in near-real time”, which is what the number now says.
- Activity is not normalised across sources. The score is a sum over incommensurable units — a machine-coded news record, a UCDP battle record and an OFAC designation each contribute through the same two weight tables. A consequence worth stating plainly: adding or removing a source rescales every theater it touches, retroactively.
- Hazards can be double-counted across sources. USGS and GDACS may both report the same earthquake, and because event ids are content-addressed per source they are two events, each contributing separately. Cross-source hazard de-duplication is not built.
- News-derived events are tripwires, not findings — but they still add up. GDELT-sourced candidates enter at low confidence, are placed by the coordinates in the coded record, and are never re-labelled as confirmed. They are not, however, prevented from carrying a score on their own: because activity is an unnormalised sum, a theater covered mainly by machine-coded news reaches a high band on recent volume, and on the live feed several do. Read those bands with the source composition beside them.
- Two rules can promote a machine-coded candidate to medium — and never past it. Both double that row’s confidence weight (0.3 → 0.6) and both are tagged on the event, so a promotion is auditable rather than invisible. Primary-source path: a high-confidence event from a different source shares the candidate’s theater, event type and date (
corroborated-by:<source>). It is a partial double count — the corroborating primary event is already in the feed carrying its own weight — and it fires rarely, a handful of rows in a 120-day window. Breadth path (methodology 2026.09.1): a machine-coded row carried by at least 3 distinct outlets (corroboration, kept through the dedup that used to discard it, plus the day’s mentions of the same coded event) becomes medium (corroborated-by:breadth:<n>). This one is common — roughly a third of a day’s machine-coded rows on the two replayed days — and it is why the Bab al-Mandab island seizure of 2026-09-10 now grades medium instead of low. Neither path reaches high: sixty outlets repeating one wire story is still news, and only a primary or official source is graded high, by the adapter that read it. - Machine-coded rows are now gated on the coder’s own confidence, and the gate counts what it drops. GDELT publishes a per-article extraction confidence (10–100). A coded event whose best article scores below 30 is not published; the count is printed in the daily health readout rather than folded into a smaller number. The row’s link is the highest-confidence article (arepresentative article, labelled so), not whichever row happened to carry the widest mention count. A row whose mentions file was missing that quarter-hour is kept with its confidence marked as not measured — the events are never hostage to the companion file.
- An organization’s analyst can overrule a grade — for that organization only. A member an owner has designated as an analyst may set an event’s confidence, theater or type for their organization, with a required reason. The shared feed row is never edited: the override is applied where that organization reads the event — its register, feed, brief, alerts and stored snapshots — and is labelled analyst-adjusted beside the pipeline’s value wherever it appears, with the reason one hover away. No other organization sees it, and this page’s numbers never include one. Each override also joins the field-report log as calibration evidence; whether it changes the shared feed is a separate, published decision.
Review us — candor, not encouragement
This instrument is young and we would rather hear its weaknesses from a reviewer than from a customer. If you teach or research international security, business continuity, or risk measurement, we invite a red-pen read of four things, in whatever depth you have time for:
- The event taxonomy — are the types observable, mutually distinguishable, and useful to a continuity planner? What is missing, what should be split, what should be merged?
- The confidence and likelihood method above — where does the weighting mislead, and what would a defensible recalibration look like once history accumulates?
- The exercise (wargame) doctrine — do the scenario ladders in the demo reflect how coercion actually escalates, and what would you add or strike?
- The framing — anywhere the product implies more certainty than the evidence supports.
What we commit to in return: substantive critiques are answered in writing, and changes they cause are recorded against the methodology version below with attribution if you want it. We keep a public corrections record rather than a quiet edit history. Send reviews to bhyde@engsecsolutions.com.
Citing this
This methodology is versioned: argus-methodology-2026.10. The constants on this page are imported live from the engine, which makes the page self-updating and therefore a moving target — a number quoted from it in March is not necessarily the number it renders in June. Cite the version alongside the claim so the two can be told apart:
Argus geopolitical risk register, methodology argus-methodology-2026.10, retrieved <date>.
The version is bumped whenever weights, thresholds or the taxonomy change, and a posture comparison across a bump compares two different rulers — the product says so out loud rather than letting the line move silently. Version history, the calibration plan and the corrections record are on the track record page. What is not yet available is a dated, immutable snapshot of the feed itself: exposure is recomputed over a rolling 120-day window, so a register value is reproducible only for as long as the window still holds its evidence. Until snapshots ship, treat a cited score as an observation made on a date, not a re-runnable query. Source-level terms — including the ones that require attribution when you republish — are on sources & licensing.
Frameworks & references
The register follows TARA-style threat-asset-risk assessment; the asset/dependency model maps to an ISO 22301 Business Impact Analysis (§8.2.2) and risk assessment (§8.2.3); treatments follow ISO 31000 §6.5 decisions (accept · mitigate · transfer · avoid) with residual scoring. The DIME instrument dimension (Diplomatic · Informational · Military · Economic) follows the instruments-of-national-power framing in U.S. joint doctrine (JP 1). Framework mappings in the product are worded as evidence for — never “makes you compliant”.