VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

The False Positive Laboratory | Where the Forecast Was Worried and Nothing Happened | The Purple Report

This article is part of The Purple Report September 2026 | Disaster Forecasting and Predictions and focuses on how forecasts are checked against later outcomes.

A forecasting project can look brilliant if it remembers only the disasters it anticipated.

That is not enough.

We also need to remember every place we worried about where nothing happened.

Every place we failed to worry about where something did happen.

Every correct hotspot with the wrong mechanism.

Every warning that appeared to be a false alarm because preparation successfully prevented catastrophe.

A forecasting system earns trust by preserving its mistakes at the same resolution as its successes.

Quick Read

The Purple Forecast will keep an immutable monthly baseline and classify later outcomes as:

  • HIT: a declared hazard or state materialised inside the declared scope.
  • FALSE POSITIVE: a time-bounded elevated concern expired without the expected event.
  • MISS: a material event occurred without the forecast identifying the system appropriately.
  • CORRECT NEGATIVE: a time-bounded forecast indicated low concern and the event did not occur.
  • WRONG MECHANISM: the hotspot was right but the event occurred through a different process.
  • RIGHT HAZARD, WRONG CONSEQUENCE: the physical event occurred but the consequence path differed.
  • DATA GAP: a defensible judgment was not possible because critical observation was unavailable.
  • SUCCESSFUL PREVENTION: the hazard occurred or warning was justified, but preparation broke the expected catastrophe chain.

The purpose is not to optimise one impressive accuracy percentage.

It is to understand how the system is wrong.

Why False Positives Are Necessary

A system designed to issue concern only when catastrophe is nearly certain will miss the period when prevention is easiest.

Consider a slope under prolonged rainfall.

If the system waits until the slope is visibly collapsing before raising concern, it may achieve a low false-alarm rate and provide little useful lead time.

If it raises concern whenever a credible convergence appears, some of those convergences will dissipate without failure.

That is not automatically a bad forecast.

The question becomes:

How many false alarms are produced for each useful hit, at what lead time, and at what cost to the people and institutions acting on the forecast?

Forecast Verification Is a Discipline, Not a Victory Lap

The World Meteorological Organization’s forecast-verification guidance treats verification as a way to discover accuracy, bias, reliability, resolution and uncertainty—not merely a ceremonial score reported after a forecast.

Hydrological forecast guidance emphasises several different qualities:

  • accuracy — how close forecasts are to outcomes;
  • bias — whether a system systematically over- or under-forecasts;
  • reliability — whether forecast probabilities match observed frequencies;
  • resolution — whether the forecast can distinguish more dangerous from less dangerous states;
  • sharpness — how much useful confidence the system can express.

No single measure contains all of these.

WMO — Guidelines for Streamflow Forecast Verification

The Four-Box Test

For time-bounded yes/no forecast questions, the basic logic is simple:

Event occurredEvent did not occur
Concern issuedHitFalse alarm
Concern not issuedMissCorrect negative

Operational meteorology has long used measures such as probability of detection, false-alarm ratio, bias and critical-success scores to analyse this table.

The Purple Report will use the logic without pretending every hazard should optimise the same numerical target.

A tsunami warning and a ten-year structural subsidence watch have completely different costs, lead times and event definitions.

Probability Forecasts Need Reliability

If a forecast repeatedly says “70% chance”, then over a sufficiently large comparable set, the event should occur roughly seven times out of ten if the forecast is well calibrated.

That is what reliability means conceptually.

A system that says 90% whenever it is nervous may sound authoritative and prove badly calibrated.

This is one reason the September baseline avoids manufacturing exact percentages where authorities have not provided a defensible model.

WATCH, ELEVATED and HIGH-CONCERN are research states.

They are not disguised probabilities.

The Gyirong Test Must Be Blind

Gyirong is dangerous for this project because we already know the answer.

After a catastrophe, almost every prior observation can be made to look like a warning sign.

That is hindsight bias.

The retrospective experiment therefore freezes the clock at multiple pre-event dates and hides all future information.

The model must rank Gyirong using only what existed at that moment.

Then the same ranking process is run against comparable Himalayan corridors where nothing catastrophic occurred.

That comparison is essential.

A model that identifies every Himalayan valley as HIGH-CONCERN has not predicted Gyirong. It has coloured a mountain range red.

Wrong Mechanism Is an Important Result

Suppose the Purple Forecast correctly identifies a mountain corridor as unusually dangerous because of glacial-lake and landslide exposure.

Six months later, an earthquake triggers the major slope failure instead of the anticipated rainfall or glacier mechanism.

Was the forecast correct?

Partly.

The forecast identified the exposed people, infrastructure and structural vulnerability correctly.

It misunderstood the trigger.

That should be recorded as WRONG MECHANISM, not quietly celebrated as a hit.

This distinction teaches causal structure rather than just location correlation.

Right Hazard, Wrong Consequence

Suppose a forecast anticipates a severe cyclone and expects long-duration national power loss.

The cyclone makes landfall, but strong grid redundancy and pre-positioned repair teams restore power rapidly.

The hazard forecast was right.

The cascade forecast was wrong.

That is valuable.

The model has discovered resilience.

It should lower future expected cascade severity when the same verified repair capacity remains available.

Successful Prevention Is Not a False Alarm

This is one of the hardest verification problems.

A warning is issued.

People evacuate.

The hazard arrives.

Few people die because the evacuation worked.

If we verify only against catastrophe outcome, the warning might appear exaggerated.

But prevention changed the outcome.

The forecast therefore separates:

  • physical hazard occurrence;
  • warning correctness;
  • protective action;
  • final consequence.

The absence of catastrophe after successful preparation is evidence of repair capacity, not necessarily evidence that the forecast was wrong.

The September 2026 Baseline Is the Control

The September master report is deliberately preserved as a dated baseline.

Future editions should never rewrite September’s original watch states to make the past look smarter.

Instead they append:

  • what new evidence arrived;
  • which state changed;
  • why it changed;
  • what eventually happened;
  • whether the forecast was too sensitive or too quiet;
  • what should change in the next model.

How We Will Detect Alarm Bias

A disaster project can develop alarm bias when researchers spend too much time looking only for failure.

Once the researcher spends all day looking for failure, every anomaly begins to look ominous.

We therefore need explicit negative tests:

  • search for evidence that the hotspot is becoming safer;
  • search for official downgrades;
  • search for rainfall/heat/unrest forecasts weakening;
  • identify repaired infrastructure;
  • find comparable locations that remained stable;
  • ask whether the same signal appeared many times before without failure;
  • look for another explanation of the observed anomaly.

The review’s job is not to prove danger.

It is to distinguish among possible states.

The Minimum Public Scorecard

MeasureQuestion
HitsHow many declared time-bounded concerns produced the defined event?
False alarmsHow many expired without the defined event?
MissesHow many material events were not appropriately on the ledger?
Lead timeHow long before the event did useful concern emerge?
Wrong mechanismHow often was place/consequence right but trigger wrong?
State reversalsDid the system downgrade when evidence improved?
Data gapsHow many misses were caused by absent observation?
Prevention successHow often did preparation materially break the predicted cascade?

The Ten-Year Integrity Rule

By 2036, the Purple Forecast should be able to show not only the disasters it anticipated.

It should show:

  • every original 2026 hotspot;
  • every monthly state change;
  • every major miss;
  • every false alarm;
  • every important data gap;
  • every model correction;
  • every case where resilience stopped the expected cascade.

If the public record contains only our successes, the forecast has failed its most important verification test.

Primary Method Sources