A test is 95% accurate.
You test positive.
How likely is it that you actually have the condition?
Many people instinctively answer: about 95%.
But that answer can be wildly wrong.
If the condition is extremely rare, most positive results can still come from false positives simply because the healthy population is enormous.
The missing piece is the base rate: how common the condition was before the new test result arrived.
When people focus on the vivid new clue and neglect the background prevalence, they commit base-rate neglect.
Quick Read
Base-rate neglect is the tendency to underweight background prevalence when judging a specific case.
The base rate answers a simple question:
Before I learned anything special about this case, how common was this outcome in the relevant population?
Case-specific evidence then changes that prior probability.
Good reasoning uses both.
Base-rate neglect occurs when the case description, stereotype, test result or story is allowed to overwhelm the statistical background.
This mechanism became famous through research on judgment under uncertainty associated with Daniel Kahneman and Amos Tversky, especially work on representativeness and probabilistic reasoning.
The One-Sentence Answer
Base-rate neglect works when a vivid or seemingly diagnostic description of an individual case receives more weight than the background frequency needed to interpret that evidence correctly.
The Base-Rate Chain
population prevalence → prior probability → new case evidence → likelihood of evidence under competing states → updated probability → decision
Base-rate neglect breaks the chain near the beginning.
The prior is forgotten, dismissed or given too little weight.
Base Rate Is Not Destiny
A base rate does not tell you what will happen to one specific person.
It tells you where the inference should begin.
If only 1% of applicants in a broad pool become elite performers, that does not mean a particular candidate has a 1% chance after you learn they have exceptional evidence.
New evidence should update the base rate.
The mistake is either ignoring the prior completely or refusing to update it at all.
Base-Rate Neglect Is Not Bayesian Inference
Bayesian inference is the formal machinery for combining prior probability with new evidence.
Base-rate neglect is one way human judgment can fail to perform that update appropriately.
eduKateSG already has a canonical article for the updating machinery: How Bayesian Inference Works.
This article owns the behavioural failure mode: the prior is available, relevant and still underweighted.
Base-Rate Neglect Is Not Availability
Availability asks what comes to mind easily.
Base-rate neglect asks whether prevalence is properly included in the inference.
A vivid anecdote can be highly available and therefore crowd out a base rate.
But the two mechanisms remain distinct.
See How The World Works | Availability Heuristic.
Representativeness: The Story That Looks Right
Suppose you read this description:
Alex is quiet, precise, enjoys abstract problems and spends weekends reading technical manuals.
Is Alex more likely to be an engineer or a salesperson?
The description may feel representative of an engineer stereotype.
But if there are twenty times more salespeople than engineers in the population from which Alex was drawn, that base-rate information matters.
A resemblance-based story can feel diagnostic even when the population structure points elsewhere.
The Medical Screening Example
Imagine a condition affects 1 in 1,000 people.
A test catches 99% of true cases and has a 1% false-positive rate among people without the condition.
Now imagine 100,000 people are tested.
- About 100 people truly have the condition.
- About 99 of them test positive.
- About 99,900 people do not have it.
- About 999 of those healthy people test positive falsely.
So there are roughly 1,098 positive tests, but only about 99 true cases among them.
The positive test is important.
It is not interpretable without prevalence.
Natural Frequencies Help
People often reason more clearly when probabilities are translated into counts.
Instead of:
“Prevalence 0.1%, sensitivity 99%, false-positive rate 1%,”
say:
“Out of 100,000 people, about 100 have the condition. Roughly 99 of those test positive, while about 999 healthy people also test positive.”
The denominator becomes visible.
The Hiring Example
A candidate gives an extraordinary interview.
The panel leaves impressed.
But how predictive are interviews of later performance in this role?
How often do candidates with similar profiles succeed?
What is the historical base rate of strong performance after a top interview score?
The interview is case evidence.
Historical performance is base-rate evidence.
Good hiring combines them.
The Startup Example
The founder is charismatic.
The product demo is excellent.
The market story is compelling.
None of those facts should be ignored.
But venture outcomes have base rates.
Revenue milestones have base rates.
Expansion plans have base rates.
The stronger the narrative, the more important it becomes to restore the reference class.
The Investment Example
An analyst discovers a company with a brilliant product and rapidly growing sales.
The temptation is to forecast from the company story alone.
But how often do companies at this stage sustain this growth?
How often do margins compress?
How often does competition arrive?
The outside view does not invalidate company-specific evidence.
It calibrates it.
The Criminal-Justice Example
Risk assessment can go wrong when vivid case details receive weight without properly calibrated population statistics.
But the opposite error is also dangerous: treating group statistics as destiny for an individual.
Base rates belong inside probabilistic reasoning, not as substitutes for individual evidence, rights or due process.
High-stakes systems need explicit safeguards because statistical relevance and ethical legitimacy are different questions.
Base Rates Depend on the Right Population
“What is the base rate?” is incomplete.
The correct question is:
What is the base rate in the population relevant to this case?
A national rate may be too broad.
A local subgroup may be more appropriate.
But narrowing the population too aggressively can overfit the case.
Reference-class selection is itself a modelling decision.
The Reference-Class Problem
A construction project asks how long it will take.
Should we compare with all construction projects?
Bridges?
Bridges of this size?
Bridges in this regulatory environment?
The base rate improves as the class becomes more relevant—but can become unstable if the class becomes too small.
Good judgment balances relevance and sample size.
Base Rates Can Be Bad Data
Sometimes people discount base rates because the rates are poor.
The population is outdated.
The measurement is weak.
The context has changed.
The case belongs to a different subgroup.
Modern research on base-rate use is therefore more nuanced than “people irrationally ignore statistics.”
People may rationally discount a base rate they perceive as untrustworthy or irrelevant.
The audit must test the base rate before insisting on it.
Case Evidence Can Be Extremely Strong
A rare disease has a low base rate.
A highly specific genetic test may still make the posterior probability very high.
A rare machine fault has a low base rate.
A distinctive failure code may still be powerful evidence.
Respecting base rates does not mean worshipping them.
It means weighting the prior and the likelihood together.
Likelihood Ratios: How Diagnostic Is the Evidence?
The right question about a clue is not merely “Does this clue occur when the condition is present?”
Ask:
How much more likely is this evidence if the condition is present than if it is absent?
A clue common in both states has weak diagnostic value even if it sounds relevant.
This is the bridge from storytelling to likelihood.
Base-Rate Neglect and Confirmation Bias
A person already believes a hypothesis.
They notice one vivid confirming case.
The case receives enormous weight.
The low base rate of such cases in the wider population is ignored.
Now two biases reinforce one another.
Confirmation bias selects the convenient evidence.
Base-rate neglect misweights it.
The Education Example
A parent hears about one student who improved thirty marks in six weeks.
The case is real.
Should every student expect the same?
Not without asking how often such gains occur among students with similar starting scores, attendance, syllabus gaps, time available and examination proximity.
The individual success story proves possibility.
The base rate estimates typicality.
Diagnostic Testing in School
A screening quiz flags a student as “at risk.”
The flag should not be treated as a diagnosis by itself.
How common is the underlying difficulty in this cohort?
How sensitive is the screen?
How specific?
What other evidence exists?
Screening should route investigation, not replace it.
Base Rates and Forecasting
Forecasts often become optimistic because teams focus on the unique details of the current plan.
The base-rate correction is the outside view:
How long did comparable projects actually take?
How often did costs exceed budget?
How often did demand miss the forecast?
This becomes central in the Planning Fallacy article later in this batch.
Base Rates and Probability
A conditional probability without a base rate can be misleading.
“90% of people with the condition show symptom X” does not tell you the probability of the condition given symptom X.
You also need to know how often symptom X appears among people without the condition and how common the condition is overall.
Confusing P(evidence | condition) with P(condition | evidence) is a common diagnostic error.
The Prosecutor’s Fallacy
Suppose a forensic match is very rare among innocent people.
It is a mistake to jump directly from “rare match if innocent” to “therefore almost certainly guilty.”
The prior probability, alternative explanations, search process and total number of possible matches matter.
High-stakes evidence requires explicit conditional reasoning because intuitive inversion can be dangerous.
Search Changes the Base Rate
If you search one person for a rare pattern, finding it can be strong evidence.
If you search one million people for the same rare pattern, some matches may arise by chance.
The process that generated the candidate matters.
This connects base-rate reasoning to multiple testing and selection.
Base Rates Change Over Time
A historical rate can become stale.
Technology changes.
Vaccination changes prevalence.
New regulations change failure rates.
Education systems change syllabus.
The prior should reflect the current generating process, not whatever statistic is easiest to find.
Base Rates and Legibility
Large systems create categories so base rates can be calculated.
Failure rate by machine type.
Readmission rate by condition.
Examination outcome by starting band.
But category design matters.
A bad classification system produces misleading priors.
See How The World Works | Legibility.
Base Rates and Information Asymmetry
One side may know the base rate while the other sees only the case.
An insurer sees thousands of claims.
The customer sees one household.
A platform sees millions of transactions.
The seller sees one sale.
Population-level information can therefore create institutional advantage.
See How The World Works | Information Asymmetry.
The Base-Rate Audit
- Define the target outcome. What are you trying to estimate?
- Choose the relevant population. Which reference class genuinely contains this case?
- Find the prevalence. How common is the outcome before case evidence?
- Check data quality. Is the base rate current, measured well and contextually relevant?
- Write the case evidence separately. What new information does this case add?
- Ask how diagnostic it is. How often would this clue appear under each competing state?
- Use natural frequencies. Convert percentages to expected counts where helpful.
- Do not invert conditional probabilities. P(A|B) is not automatically P(B|A).
- Check availability. Is one vivid story crowding out the denominator?
- Check representativeness. Does the case merely resemble a stereotype?
- Compare several reference classes. Is the conclusion robust to reasonable population definitions?
- Update rather than replace. New evidence should modify the prior, not erase it automatically.
- Check for regime change. Has the historical base rate become stale?
- Document the posterior logic. What changed and why?
When the Base-Rate Lens Fails
The lens fails when background statistics are treated as more authoritative than highly diagnostic current evidence.
It fails when the wrong reference class is chosen.
It fails when outdated or biased population data are treated as neutral priors.
And it fails when statistical group rates are used to erase individual evidence, rights or context in high-stakes decisions.
A Better Question Than “Does This Story Fit?”
Ask:
Before I heard this compelling story, how common was the outcome—and how much should this specific evidence rationally move me away from that starting point?
How Base-Rate Neglect Connects to the Rest of the World
- Probability: base rates are prior probabilities in conditional reasoning.
- Bayesian inference: formal updating combines priors with diagnostic evidence.
- Availability: vivid cases can crowd out background prevalence.
- Sampling: base rates depend on correctly defined and measured populations.
- Sorting: the observed subgroup may differ systematically from the wider population.
- Information asymmetry: institutions often know population rates that individuals do not.
- Forecasting: reference-class outcomes discipline inside-view stories.
- Legibility: categories determine which base rates can be calculated.
- Confirmation bias: preferred case evidence can be overweighted while inconvenient priors are ignored.
- Planning fallacy: project teams neglect completion base rates of comparable projects.
Frequently Asked Questions
What is a base rate?
A base rate is the background prevalence or frequency of an outcome in the relevant population before case-specific evidence is considered.
What is base-rate neglect?
It is the tendency to underweight relevant background prevalence when judging an individual case, often because vivid or representative case evidence receives too much attention.
Does base-rate neglect mean we should ignore individual evidence?
No. Good inference combines the prior with case evidence. Strong evidence can move probability dramatically away from the base rate.
Why are natural frequencies useful?
Counts such as “99 true positives and 999 false positives out of 100,000 people” often make the denominator and comparison groups easier to understand than several conditional percentages.
Research Basis and Further Reading
- Classic Tversky–Kahneman research on representativeness and probabilistic judgment established the importance of base-rate use and neglect.
- Bayesian inference provides the normative framework for combining prior prevalence with new evidence.
- Later research has shown that base-rate use depends on relevance, trust, presentation and the perceived diagnosticity of case evidence, making the phenomenon more context-sensitive than a simple universal rule.
What to Read Next on eduKateSG
- How Probability Works — the conditional probability machinery underneath the problem.
- How Bayesian Inference Works — how priors and likelihoods become posteriors.
- How The World Works | Availability Heuristic — why vivid examples become overweighted.
- How Forecasting Works — how historical reference classes improve future estimates.
The Larger Idea
Stories tell us what can happen.
Base rates tell us how often worlds like that usually occur.
Neither is enough alone.
A statistic without case evidence can be blind to what is special.
A story without prevalence can mistake exceptional detail for typical reality.
Base-rate neglect is what happens when the case in front of us becomes so vivid that we forget the population standing behind it.