This article is part of The Purple Report September 2026 | Disaster Forecasting and Predictions and focuses on how forecasts are checked against later outcomes.
A forecasting project can look brilliant if it remembers only the disasters it anticipated.
That is not enough.
We also need to remember every place we worried about where nothing happened.
Every place we failed to worry about where something did happen.
Every correct hotspot with the wrong mechanism.
Every warning that appeared to be a false alarm because preparation successfully prevented catastrophe.
A forecasting system earns trust by preserving its mistakes at the same resolution as its successes.
Quick Read
The Purple Forecast will keep an immutable monthly baseline and classify later outcomes as:
- HIT: a declared hazard or state materialised inside the declared scope.
- FALSE POSITIVE: a time-bounded elevated concern expired without the expected event.
- MISS: a material event occurred without the forecast identifying the system appropriately.
- CORRECT NEGATIVE: a time-bounded forecast indicated low concern and the event did not occur.
- WRONG MECHANISM: the hotspot was right but the event occurred through a different process.
- RIGHT HAZARD, WRONG CONSEQUENCE: the physical event occurred but the consequence path differed.
- DATA GAP: a defensible judgment was not possible because critical observation was unavailable.
- SUCCESSFUL PREVENTION: the hazard occurred or warning was justified, but preparation broke the expected catastrophe chain.
The purpose is not to optimise one impressive accuracy percentage.
It is to understand how the system is wrong.
Why False Positives Are Necessary
A system designed to issue concern only when catastrophe is nearly certain will miss the period when prevention is easiest.
Consider a slope under prolonged rainfall.
If the system waits until the slope is visibly collapsing before raising concern, it may achieve a low false-alarm rate and provide little useful lead time.
If it raises concern whenever a credible convergence appears, some of those convergences will dissipate without failure.
That is not automatically a bad forecast.
The question becomes:
How many false alarms are produced for each useful hit, at what lead time, and at what cost to the people and institutions acting on the forecast?
Forecast Verification Is a Discipline, Not a Victory Lap
The World Meteorological Organization’s forecast-verification guidance treats verification as a way to discover accuracy, bias, reliability, resolution and uncertainty—not merely a ceremonial score reported after a forecast.
Hydrological forecast guidance emphasises several different qualities:
- accuracy — how close forecasts are to outcomes;
- bias — whether a system systematically over- or under-forecasts;
- reliability — whether forecast probabilities match observed frequencies;
- resolution — whether the forecast can distinguish more dangerous from less dangerous states;
- sharpness — how much useful confidence the system can express.
No single measure contains all of these.
WMO — Guidelines for Streamflow Forecast Verification
The Four-Box Test
For time-bounded yes/no forecast questions, the basic logic is simple:
| Event occurred | Event did not occur | |
|---|---|---|
| Concern issued | Hit | False alarm |
| Concern not issued | Miss | Correct negative |
Operational meteorology has long used measures such as probability of detection, false-alarm ratio, bias and critical-success scores to analyse this table.
The Purple Report will use the logic without pretending every hazard should optimise the same numerical target.
A tsunami warning and a ten-year structural subsidence watch have completely different costs, lead times and event definitions.
Probability Forecasts Need Reliability
If a forecast repeatedly says “70% chance”, then over a sufficiently large comparable set, the event should occur roughly seven times out of ten if the forecast is well calibrated.
That is what reliability means conceptually.
A system that says 90% whenever it is nervous may sound authoritative and prove badly calibrated.
This is one reason the September baseline avoids manufacturing exact percentages where authorities have not provided a defensible model.
WATCH, ELEVATED and HIGH-CONCERN are research states.
They are not disguised probabilities.
The Gyirong Test Must Be Blind
Gyirong is dangerous for this project because we already know the answer.
After a catastrophe, almost every prior observation can be made to look like a warning sign.
That is hindsight bias.
The retrospective experiment therefore freezes the clock at multiple pre-event dates and hides all future information.
The model must rank Gyirong using only what existed at that moment.
Then the same ranking process is run against comparable Himalayan corridors where nothing catastrophic occurred.
That comparison is essential.
A model that identifies every Himalayan valley as HIGH-CONCERN has not predicted Gyirong. It has coloured a mountain range red.
Wrong Mechanism Is an Important Result
Suppose the Purple Forecast correctly identifies a mountain corridor as unusually dangerous because of glacial-lake and landslide exposure.
Six months later, an earthquake triggers the major slope failure instead of the anticipated rainfall or glacier mechanism.
Was the forecast correct?
Partly.
The forecast identified the exposed people, infrastructure and structural vulnerability correctly.
It misunderstood the trigger.
That should be recorded as WRONG MECHANISM, not quietly celebrated as a hit.
This distinction teaches causal structure rather than just location correlation.
Right Hazard, Wrong Consequence
Suppose a forecast anticipates a severe cyclone and expects long-duration national power loss.
The cyclone makes landfall, but strong grid redundancy and pre-positioned repair teams restore power rapidly.
The hazard forecast was right.
The cascade forecast was wrong.
That is valuable.
The model has discovered resilience.
It should lower future expected cascade severity when the same verified repair capacity remains available.
Successful Prevention Is Not a False Alarm
This is one of the hardest verification problems.
A warning is issued.
People evacuate.
The hazard arrives.
Few people die because the evacuation worked.
If we verify only against catastrophe outcome, the warning might appear exaggerated.
But prevention changed the outcome.
The forecast therefore separates:
- physical hazard occurrence;
- warning correctness;
- protective action;
- final consequence.
The absence of catastrophe after successful preparation is evidence of repair capacity, not necessarily evidence that the forecast was wrong.
The September 2026 Baseline Is the Control
The September master report is deliberately preserved as a dated baseline.
Future editions should never rewrite September’s original watch states to make the past look smarter.
Instead they append:
- what new evidence arrived;
- which state changed;
- why it changed;
- what eventually happened;
- whether the forecast was too sensitive or too quiet;
- what should change in the next model.
How We Will Detect Alarm Bias
A disaster project can develop alarm bias when researchers spend too much time looking only for failure.
Once the researcher spends all day looking for failure, every anomaly begins to look ominous.
We therefore need explicit negative tests:
- search for evidence that the hotspot is becoming safer;
- search for official downgrades;
- search for rainfall/heat/unrest forecasts weakening;
- identify repaired infrastructure;
- find comparable locations that remained stable;
- ask whether the same signal appeared many times before without failure;
- look for another explanation of the observed anomaly.
The review’s job is not to prove danger.
It is to distinguish among possible states.
The Minimum Public Scorecard
| Measure | Question |
|---|---|
| Hits | How many declared time-bounded concerns produced the defined event? |
| False alarms | How many expired without the defined event? |
| Misses | How many material events were not appropriately on the ledger? |
| Lead time | How long before the event did useful concern emerge? |
| Wrong mechanism | How often was place/consequence right but trigger wrong? |
| State reversals | Did the system downgrade when evidence improved? |
| Data gaps | How many misses were caused by absent observation? |
| Prevention success | How often did preparation materially break the predicted cascade? |
The Ten-Year Integrity Rule
By 2036, the Purple Forecast should be able to show not only the disasters it anticipated.
It should show:
- every original 2026 hotspot;
- every monthly state change;
- every major miss;
- every false alarm;
- every important data gap;
- every model correction;
- every case where resilience stopped the expected cascade.
If the public record contains only our successes, the forecast has failed its most important verification test.