VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Intelligence Works | Reference Class Selection — How Intelligence Chooses Which Past Cases Are Similar Enough to Set the Base Rate

HOW INTELLIGENCE WORKS · REFERENCE CLASS SELECTION · eduKateSG

How Intelligence Chooses Which Past Cases Are Similar Enough to Set the Base Rate

Reference class selection is the intelligence process that decides which past cases belong in the comparison set used to estimate what usually happens. The choice matters because a base rate is only as relevant as the population from which it was drawn.

Target case → candidate comparison classes → similarity dimensions → sample size → relevance → stability → chosen class → base rate → case-specific update.

This article belongs to the How Intelligence Works series. Base-Rate Reasoning owns using background frequencies before judging the case in front of us. Reference class selection owns the upstream problem: which cases count as the right background?

There Is Usually More Than One Base Rate Available

A project can be compared with all projects, all software projects, all software projects of similar size, all projects in the same organisation, or only projects using the same technology. Each class produces a different historical rate.

Choosing the reference class is therefore not a trivial pre-processing step. It is part of the inference.

The outside view begins by deciding which outside cases are actually relevant.

1. A Reference Class Is a Chosen Comparison Population

The target case has many attributes. The reference class selects some of those attributes as relevant enough to define comparable cases.

The difficulty is that similarity can be defined in many ways, and not every similarity matters for the outcome being predicted.

2. Too Broad and Too Narrow Are Both Dangerous

Reference classStrengthRisk
Very broadLarge sample and stable rateMay mix cases governed by different mechanisms
Moderately specificBalances relevance and sample sizeStill depends on choosing the right similarity dimensions
Very narrowLooks highly tailored to the caseSmall sample, unstable rate and cherry-picking risk
ConvenientEasy data accessMay be selected because it supports the preferred forecast
Mechanism-basedGroups cases with similar causal structureMechanism may itself be uncertain
Outcome-matchedCan appear preciseMay leak knowledge of the result into the selection

3. Similarity Must Be Relevant to the Outcome

Two cases can look similar on visible features and differ on the variable that actually controls the outcome. Conversely, cases from different surface domains may share the same causal structure.

Good reference-class selection asks which attributes would reasonably change the base rate and which are merely decorative similarities.

4. The Reference Class Problem Has No Automatic Answer

A single target often belongs to many nested and overlapping classes. There may be no uniquely correct comparison set available from the description alone.

Intelligence therefore compares several plausible classes, inspects whether their estimates agree, and makes the selection rule visible.

If the forecast changes dramatically when the comparison class changes slightly, the reference-class choice is part of the uncertainty and should be reported.

5. Reference Class Selection in Mathematics and Statistics

Statistical estimates depend on the population from which cases are treated as exchangeable or comparable. A rate estimated from one population may not transfer when the generating conditions differ.

The reasoning task is not only to calculate the frequency correctly, but to decide whether the cases used to produce it belong in the same inferential population as the target.

6. Reference Class Selection in Learning

A teacher asking whether a learner is “behind” needs a comparison group. Same age? Same curriculum exposure? Same language background? Same prior achievement? Same instructional history?

Different classes answer different questions. The comparison should match the decision being made rather than supply a convenient label.

Good educational diagnosis therefore keeps the reference class visible instead of treating the resulting percentile or rate as context-free truth.

7. Forecasting Needs an Outside View Before the Inside Story Takes Over

Plans feel unique from the inside because their details are vivid. Reference classes provide a corrective by asking how similar efforts actually turned out.

But the outside view becomes useful only after the comparison set is chosen responsibly. A project team can always find a flattering class if the selection rule is allowed to move after the desired forecast is known.

8. Multiple Reference Classes Can Be a Signal, Not a Nuisance

If several defensible classes produce similar base rates, confidence in the outside view increases. If they diverge, the disagreement identifies structural uncertainty.

Instead of hiding that divergence, intelligence can report a range or weight classes according to relevance and evidence quality.

9. Reference-Class Failure Atlas

FailureWhat happensRepair
Broad-class dilutionUnrelated cases wash out relevant structureNarrow by outcome-relevant mechanism
Narrow-class instabilityTiny sample gives a volatile rateBroaden until precision is usable
Convenience samplingAvailable data substitutes for relevant dataDefine the class before retrieval
Cherry-picked classComparison set is chosen to support the preferred answerPrecommit selection criteria
Surface matchingVisible similarity hides different causal regimesCompare mechanisms
Outcome leakageKnowledge of the result influences class membershipSelect without using outcome information
Single-class certaintyAlternative defensible classes are ignoredRun sensitivity across classes

10. Reference Class Selection and Base-Rate Reasoning Are Different

Reference class selection defines the population from which a base rate is estimated. Base-rate reasoning decides how that background rate should influence judgement of the present case.

One chooses the denominator. The other uses the denominator.

11. Teams Should Argue About the Class Before Arguing About the Forecast

Many forecast disputes are really reference-class disputes in disguise. One analyst compares the target with all prior projects; another compares it only with recent projects using the same technology.

Make the comparison class explicit before debating the number it produces.

12. Institutions Need Reference-Class Governance

Repeated decisions become more reliable when institutions preserve comparable cases, define inclusion rules and record when the operating regime changed enough to make older cases less relevant.

A historical database without a class-selection policy can create false precision because every stored case appears equally comparable.

13. Artificial Intelligence and Retrieval-Based Reference Classes

AI systems often retrieve “similar” examples before making a recommendation or prediction. The hidden question is similar in what way?

Nearest-looking cases may share wording while differing on the mechanism that matters. Reliable systems should expose the attributes used to define similarity and test whether predictions are sensitive to alternative comparison sets.

Similarity search becomes reasoning only when relevance to the target outcome is examined.

14. The Reference Class Selection Audit

  • Target: What case are we trying to judge?
  • Outcome: What quantity or event are we predicting?
  • Candidate classes: Which comparison populations are defensible?
  • Mechanism: Which similarities should affect the outcome?
  • Sample size: Is the class large enough for a stable rate?
  • Regime: Have conditions changed enough to make older cases misleading?
  • Selection rule: Was the class defined before seeing the desired answer?
  • Sensitivity: How much does the estimate move across plausible classes?
  • Transparency: Can another person reproduce the class?
  • Use: How will the chosen base rate combine with case-specific evidence?

15. CivDJ Reading: Choose Which Past Rooms Count as the Same Kind of Room

In the CivDJ frame, the Warehouse may contain thousands of prior cases. Reference class selection decides which of them are similar enough in Receiver State, constraints and mechanism to inform the current mix.

The wrong class produces a confident return from irrelevant history.

Before asking what usually happened, decide what “usually” is allowed to include.

16. Return to the Denominator

Reference class selection shows that background statistics are not context-free objects waiting to be applied.

They are produced by a choice about which cases belong together, which similarities matter and which historical regimes remain relevant.

The mature mind does not ask only, “What is the base rate?” It asks, “Base rate among which cases—and why are those the right cases for this decision?”


How Intelligence Works | Main Hub

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading