One-sentence answer: Counterfactual reasoning works by asking what would have happened to the same target under a relevant alternative condition, then using experiments, comparison groups, natural variation or explicit models to estimate that unobserved alternative without pretending it was directly observed.
“It improved after we changed X, therefore X caused the improvement” is one of the most tempting errors in reasoning. The missing question is: What would have happened if X had not changed?
That alternative outcome is usually impossible to observe for the same person, machine, school, city or system at the same moment. We see the path that happened. We do not simultaneously see the path that did not happen. Counterfactual reasoning is the disciplined attempt to bridge that missing world.
Causal claims become stronger when the alternative world is constructed by evidence and design rather than convenience.
Quick Read: the causal chain
CAUSAL QUESTION → TARGET → OBSERVED EXPOSURE / ACTION → OBSERVED OUTCOME → MISSING ALTERNATIVE OUTCOME → IDENTIFICATION STRATEGY → COMPARISON / EXPERIMENT / MODEL → ASSUMPTIONS → ESTIMATED COUNTERFACTUAL → EFFECT CONTRAST → UNCERTAINTY → SENSITIVITY TEST → DECISION BOUNDARY → NEW EVIDENCE → CORRECTION
1. The fundamental problem: we cannot observe both worlds
Suppose a student receives six weeks of tuition and improves by 12 marks. We observe:
- the student received tuition;
- the student later scored 12 marks higher.
We do not observe what the same student would have scored at the same later date without that tuition. Perhaps the student would have improved by eight marks through school lessons, maturation and independent revision. Perhaps performance would have fallen. The causal effect is not the observed 12-mark change; it is the difference between the observed outcome and the missing alternative outcome.
This is the core counterfactual problem.
2. Before-and-after is a comparison, not automatically a causal effect
A before-and-after comparison tells us that something changed over time. It does not, by itself, tell us why.
Other changes may have occurred between the two observations:
- practice or maturation;
- seasonal effects;
- economic conditions;
- a curriculum change;
- another intervention;
- selection into the programme;
- regression toward a typical value after an unusual starting point;
- changes in how the outcome was measured.
How Comparison Works therefore comes before causal attribution. Counterfactual reasoning asks whether the comparison is strong enough to stand in for the missing alternative world.
3. Potential outcomes make the missing world explicit
One influential causal framework represents each unit as having potential outcomes under different treatments or conditions. We observe the outcome corresponding to the condition that actually occurred; the other potential outcome remains counterfactual.
This formalism matters because it forces the analyst to name the missing quantity instead of silently substituting a convenient comparison. It turns “Did the programme work?” into something closer to:
For this defined population and outcome, over this time window, how would outcomes differ under programme participation versus the relevant alternative condition?
That is a testable research design problem rather than a rhetorical claim.
4. Randomisation is powerful because it builds the alternative prospectively
In a well-executed randomised experiment, assignment to treatment and control is determined by chance. Across sufficiently large groups, this helps balance both measured and unmeasured pre-treatment characteristics in expectation, allowing the control group to estimate what would have happened to the treated group without the intervention.
Randomisation does not make an experiment automatically perfect. Attrition, non-compliance, measurement problems, spillovers, small samples, poor implementation and limited external validity can still matter. But it creates a particularly transparent bridge to the counterfactual because the assignment mechanism is known.
5. When randomisation is unavailable, assumptions do more work
Many important questions cannot be answered by random assignment. We cannot randomly assign people to earthquakes, long-term poverty, many diseases, historical events or national policies. Analysts therefore use observational and quasi-experimental strategies to approximate the missing alternative.
| Strategy | Counterfactual idea | Key vulnerability |
|---|---|---|
| Matching / weighting | Compare cases similar on observed pre-treatment characteristics. | Unmeasured confounding can remain. |
| Difference-in-differences | Compare changes over time between treated and comparison groups. | Requires a credible untreated trend assumption. |
| Regression discontinuity | Compare cases near a rule-based threshold. | Effect is usually local to the threshold and design assumptions matter. |
| Instrumental variables | Use variation that shifts exposure without directly shifting the outcome except through that exposure. | Instrument validity is demanding and often contestable. |
| Synthetic control | Construct a weighted comparison from other units. | Depends on donor pool and pre-intervention fit. |
| Structural / causal model | Represent causal relationships explicitly and simulate alternatives. | Conclusions inherit model assumptions. |
OECD and European Commission work on counterfactual impact evaluation stresses exactly this point: credible evaluation depends on high-quality data, suitable comparison strategies and institutional capability, not merely the existence of a before-and-after dataset.
6. Identification is the bridge from observed data to the missing outcome
A causal effect is identified when the design and assumptions are sufficient to connect the observed data to the causal quantity being claimed. Identification is conceptually different from estimating the number precisely.
You can have:
- a very precise estimate of the wrong causal quantity;
- a correctly identified effect estimated with substantial uncertainty;
- a strong descriptive pattern for which the causal effect is not identified at all.
World-class reasoning therefore asks first, “What assumptions make this comparison causal?” and only then, “What is the estimated effect?”
7. Counterfactual ≠ fantasy scenario
A counterfactual is not simply any imaginable alternative. For causal inference, it must be connected to the observed world by a defensible design or model.
| Scenario | “Suppose demand doubled next year.” A conditional future used for exploration. |
|---|---|
| Counterfactual | “What would this unit’s outcome have been under the relevant alternative exposure?” A missing causal outcome. |
| Forecast | “What is expected to happen next year?” A time-bounded expectation. |
| Simulation | Execution of a model under specified conditions; it may help explore scenarios or counterfactuals but does not make them empirically valid by itself. |
This distinction protects causal language from becoming unfalsifiable storytelling.
8. Causal diagrams make assumptions visible
Causal diagrams can represent hypothesised relationships among exposure, outcome, confounders, mediators and selection mechanisms. Their value is not decorative. They force the analyst to declare a causal structure that can be criticised.
A variable that should be controlled in one causal structure may introduce bias in another. Blindly “controlling for everything” is therefore not a universal solution. The graph, design and causal question must agree.
9. Sensitivity analysis tests the assumptions we cannot fully verify
Some causal assumptions cannot be proven from the observed data alone. Sensitivity analysis asks how strong a violation would need to be to materially change the conclusion.
Examples include:
- How strong would an unmeasured confounder need to be to erase the estimated effect?
- Does the result survive different matching choices?
- Does it survive alternative time windows?
- Does a placebo intervention produce a similar “effect” when no causal effect should exist?
- Do pre-intervention trends support the chosen design?
- Does a negative-control outcome reveal hidden bias?
Counterfactual reasoning becomes more credible when its weak points are actively attacked rather than hidden.
10. Worked example: did a new revision routine cause improvement?
A student changes revision routine in August and scores 10 marks higher in September. Several explanations remain possible:
- the new revision routine helped;
- the September paper was easier;
- the tested topics matched recent school teaching;
- the student had more study time;
- ordinary variation produced part of the change;
- the August score was unusually low.
A practical educational decision need not always wait for a formal trial, but the reasoning should remain calibrated. The tutor could compare performance on matched question types, use repeated measures, inspect transfer to unfamiliar items, compare error categories and test whether improvement persists. The conclusion might become “the new routine is a plausible contributor and is currently worth continuing” rather than “the routine caused exactly 10 marks of improvement.”
That is stronger reasoning because the language matches the evidence.
11. Worked example: a factory repair
A production line has a high defect rate. Engineers replace a worn component and defects fall sharply. The timing supports the repair hypothesis, but a disciplined counterfactual check asks what else changed: material batch, operator, temperature, inspection rule, machine speed or upstream process.
Strong evidence might include the failure mechanism, inspection of the removed component, replication after similar repairs, process measurements, an interrupted time pattern or a controlled test. The causal conclusion becomes stronger as alternative explanations lose plausibility.
12. Counterfactuals across domains
| Domain | Counterfactual question | Boundary |
|---|---|---|
| Medicine | What would the outcome have been under another treatment or no treatment? | Requires clinical evidence; population effects do not prescribe for an individual. |
| Education | What would learning have been without this teaching intervention? | Scores, exposure and learner differences require careful comparison. |
| Policy | What would have happened without the policy? | Macro changes and selection can confound before/after comparisons. |
| Engineering | Would the failure have occurred without this component state or design choice? | Mechanism and operating conditions matter. |
| History | How dependent was an outcome on a particular event? | Historical counterfactuals can illuminate dependence but cannot recreate controlled experiments. |
| AI | Would the decision/output have differed if a feature or intervention changed? | Model counterfactuals are not automatically causal statements about the world. |
13. Individual counterfactual explanations need special care
In machine learning, a “counterfactual explanation” may describe a nearby input that would change a model’s output—for example, “if feature X were different, the model would classify this case differently.” That can help explain the model’s decision surface.
But it does not automatically mean that changing X in the real world would cause the desired outcome. The model may rely on proxies, correlated features or a structure that is not causally valid. Counterfactual about a model output ≠ causal counterfactual about reality.
14. Common counterfactual failures
| Failure | What goes wrong | Repair |
|---|---|---|
| Post hoc causation | Because B followed A, A is assumed to have caused B. | Construct a credible alternative outcome. |
| Convenient control | The comparison group differs in important pre-treatment ways. | Improve design, matching or identification strategy. |
| Hidden assumption | A strong causal claim depends on an unstated condition. | State the identification assumptions explicitly. |
| Unmeasured confounding | A third factor affects both exposure and outcome. | Use better design, additional data and sensitivity analysis. |
| Model-as-world | A simulated alternative is treated as directly observed. | Separate model output from empirical evidence. |
| Overgeneralisation | An effect identified in one population is assumed universal. | Declare target population and transfer limits. |
| Average-effect erasure | Heterogeneous effects are hidden behind one mean. | Inspect distribution and subgroup evidence where justified. |
| Authority leakage | A causal estimate is treated as automatic permission to intervene. | Keep evidence, values, consent and authority separate. |
15. Hostile test: name the missing outcome
When someone says “X caused Y,” ask:
- What is the unit or target?
- What is the observed exposure or intervention?
- What outcome was observed?
- What alternative outcome is missing?
- What design stands in for that missing outcome?
- Which assumptions make that substitution credible?
- What competing explanation remains?
- What sensitivity test could overturn the conclusion?
If the missing outcome cannot be stated, the causal claim is probably underspecified.
16. Where Counterfactuals fits in the wider How Things Work map
Counterfactuals sits downstream of Observation, Measurement and Comparison. It connects strongly to Models, Simulation, Evidence, Decision-Making and Risk.
The next discipline is uncertainty: even a well-designed causal estimate has limits, and those limits should shape what we do next.
17. What this article does not claim
- A before-and-after change is not automatically a causal effect.
- A comparison group is not automatically a valid counterfactual.
- Randomisation does not eliminate every study limitation.
- A simulated alternative world is not an observed alternative world.
- A machine-learning counterfactual explanation is not automatically a causal intervention recipe.
- A population causal estimate does not by itself determine what should be done for one individual.
- Causal evidence does not eliminate legal, ethical, clinical or institutional authority requirements.
18. Observable mastery test
You understand counterfactual reasoning when you can hear a causal claim and identify the missing alternative outcome, explain how the design estimates it, state the key identification assumptions, separate observed difference from causal effect, name at least one plausible competing explanation and propose a sensitivity or negative-control test that could weaken the conclusion.
Authoritative source corridor
- Rubin (1974): Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies — foundational potential-outcomes formulation.
- OECD & European Commission (2025): Counterfactual impact evaluations of active labour market policies — current public-policy application using linked administrative data.
- OECD (2026): Evaluating anti-fraud strategies — experimental, quasi-experimental, non-experimental and constructed-counterfactual approaches.
- National Academies: Causal Inference and Discovery in Statistics — causal questions, assumptions and statistical evidence.
Governing idea: Do not ask only what happened after the action. Ask what evidence justifies the world you are using in place of what did not happen.