Missing data works as an inferential problem because the values we failed to observe may differ systematically from the values we did observe. The correct response is therefore not simply to fill empty cells, but to prevent avoidable missingness, understand why data became missing, define the target quantity, choose an analysis whose assumptions make that target recoverable from the observed information, and test how conclusions change under plausible departures from those assumptions.
An empty spreadsheet cell looks harmless.
Nothing is written there.
So it is tempting to treat nothing as if nothing happened.
But why is the value missing?
Did a sensor fail randomly?
Did a struggling student skip the test?
Did a patient withdraw because treatment side effects were severe?
Did a survey respondent refuse to report income because their income was unusually high?
The blank cell can contain a mechanism.
The governing question: what process made these values unobserved, what does the observed data still tell us about them, and which assumptions are required to recover the scientific quantity we care about?
Quick Read
TARGET ESTIMAND → DATA COLLECTION → PREVENT MISSINGNESS → MISSINGNESS PATTERN → REASONS FOR MISSINGNESS → OBSERVED PREDICTORS OF MISSINGNESS / VALUE → ASSUMPTION ABOUT UNOBSERVED VALUES → ANALYSIS METHOD → UNCERTAINTY → SENSITIVITY TO DEPARTURES → INTERPRETATION
SPIRIT 2025 emphasises that most randomised trials experience missing outcome or covariate data, that missingness can reduce power and introduce bias, and that strategies to prevent missing data should be developed before analysis. Modern methodological work also warns that the familiar MCAR/MAR/MNAR classification is useful but not sufficient: what matters ultimately is whether the target estimand is recoverable under the assumed causal and missingness structure.
1. Missing Data Is a Property of Observation, Not Necessarily of Reality
A participant’s blood pressure still existed even if the clinic failed to record it.
A student had some level of mastery even if they were absent on test day.
The missing value is usually unknown to the researcher, not nonexistent in the world.
This distinction turns a blank cell into an inference problem.
2. Missingness Can Affect Power Even Without Bias
If 20% of outcomes disappear completely at random, fewer observations remain.
Standard errors grow.
Confidence intervals widen.
Power falls.
Missing data can therefore damage precision even when it does not distort the expected estimate.
3. Missingness Can Also Create Systematic Bias
Suppose students with the weakest performance are most likely to skip the final assessment.
The observed students are now systematically stronger than the original group.
Analysing only observed outcomes can overestimate performance.
Missingness changes who remains visible.
4. Prevention Is Usually Better Than Statistical Repair
SPIRIT 2025 explicitly recommends planning strategies to maximise complete follow-up.
Good prevention can include:
- shorter and clearer data-collection forms;
- participant reminders;
- flexible follow-up windows;
- multiple contact methods;
- backup sensors;
- staff training;
- real-time missing-field checks;
- collecting outcomes even after treatment discontinuation where ethically and operationally possible.
No imputation method can recreate information that good study operations could have observed directly.
5. Missing Outcome Data and Missing Covariates Create Different Problems
A missing final outcome removes direct information about the treatment effect.
A missing baseline covariate can prevent adjustment or prediction.
A missing exposure can change group classification.
A missing mediator can affect mechanism analysis.
The analysis should distinguish which variable is missing and why it matters to the estimand.
6. Missingness Patterns Describe Where Gaps Occur
In longitudinal data, some participants may drop out permanently after one visit.
This produces a monotone pattern.
Others may miss one visit and return later.
This produces intermittent or non-monotone missingness.
The pattern helps identify plausible mechanisms and appropriate models.
7. Missing Completely at Random Is the Strongest Simple Assumption
Under MCAR, whether a value is missing does not depend on the observed or unobserved values relevant to the analysis.
A laboratory sample destroyed by an unrelated power failure may approximate this.
Under suitable conditions, complete-case estimates can remain unbiased, though precision is lost.
MCAR is often implausible in human follow-up because absence frequently has reasons connected to participants or outcomes.
8. Missing at Random Does Not Mean “Missing Randomly” in Everyday English
MAR is a conditional assumption.
After conditioning on relevant observed information, missingness does not depend additionally on the unseen value itself.
For example, older students and students with low mid-year scores may be more likely to miss the final exam.
If those observed variables adequately explain the missingness relationship with the missing final score, an MAR-based method may be plausible.
9. Missing Not at Random Means Unobserved Values Still Matter After Conditioning
Under MNAR, missingness depends on the unseen value itself or on unobserved factors connected to it even after considering observed data.
People with the highest incomes may be especially reluctant to disclose income even after accounting for age, occupation and location.
Patients with the worst unrecorded symptom severity may skip follow-up because they are too ill to attend.
Observed data alone cannot verify such assumptions completely.
10. MCAR, MAR and MNAR Are Assumptions About Missingness Processes, Not Labels Discovered Directly From a Dataset
A statistical test can sometimes reject simple MCAR implications.
No routine test can prove MAR because MAR concerns relationships involving values we did not observe.
Missingness assumptions therefore require substantive knowledge, study operations and sensitivity analysis.
11. The MCAR/MAR/MNAR Classification Is Useful but Not Sufficient
Modern methodological work argues that with several incomplete variables, simply labelling a dataset MAR or MNAR can be too crude.
What matters is the causal relationship among substantive variables, missingness indicators and the target estimand.
Some estimands remain recoverable under missingness structures that do not fit simplistic labels cleanly.
The target question should drive the missing-data analysis.
12. Complete-Case Analysis Changes the Analysed Population
Listwise deletion removes every observation missing any variable required by the model.
This is easy.
It can be wasteful and biased.
If complete cases differ systematically from incomplete cases, the analysed sample no longer represents the original target without additional assumptions.
13. “Only 5% Missing” Is Not Automatically Safe
A small amount of missing data can create large bias if the missing values are concentrated in an extreme or influential subgroup.
Conversely, a larger proportion of missingness may be manageable when rich observed predictors explain the process well.
Percentage missing is not the whole risk-of-bias assessment.
14. Mean Imputation Makes the Dataset Look More Certain Than It Is
Replacing missing values with the observed mean feels harmless.
It reduces variance artificially.
It weakens or distorts relationships with other variables.
It treats an uncertain missing value as though it were known exactly at the mean.
Simple mean imputation is therefore rarely defensible for inferential analysis.
15. Single Regression Imputation Has the Same False-Certainty Problem
Predict a missing score from age, baseline score and other observed variables.
Insert the predicted value.
The predicted value is not actually observed.
Using one deterministic prediction ignores uncertainty about the imputed value and can underestimate standard errors.
16. Last Observation Carried Forward Freezes a Process That May Be Changing
In longitudinal studies, an old approach replaces every missing later outcome with the participant’s last observed value.
This assumes no further change after dropout.
That can be severely unrealistic in diseases, learning trajectories or recovery processes.
A convenient rule is still a model.
17. Multiple Imputation Represents Missing-Value Uncertainty With Several Completed Datasets
Multiple imputation does not guess one final value.
It generates several plausible values from an imputation model, producing multiple completed datasets.
The substantive analysis is run on each dataset.
Results are then combined so uncertainty reflects both ordinary sampling variation and uncertainty about the missing values.
18. Rubin’s Rules Combine Within- and Between-Imputation Uncertainty
Each imputed dataset yields an estimate and variance.
Rubin’s combining rules average the estimates and combine:
- uncertainty within each completed dataset;
- variation among estimates across imputations.
If the imputations disagree strongly, the final uncertainty increases.
19. The Imputation Model Should Include Variables Predictive of the Missing Value
Baseline performance can help impute a missing final score.
Prior symptom severity can help impute missing follow-up symptoms.
Including strong predictors makes MAR assumptions more plausible and improves efficiency.
Ignoring available predictive information wastes evidence.
20. Variables Predictive of Missingness Also Matter
If absence is strongly predicted by school, treatment arm or baseline health, those variables can help model the missingness mechanism.
Auxiliary variables that are not in the final scientific model can still strengthen the missing-data analysis when they predict missingness or missing values.
21. The Imputation Model Must Respect the Substantive Analysis
If the final analysis includes interactions, nonlinear terms or clustered data, a simplistic imputation model can be incompatible with it.
This is sometimes discussed as congeniality or compatibility between imputation and analysis models.
The missing-data model should preserve the relationships the substantive analysis needs to estimate.
22. Fully Conditional Specification Handles Several Variable Types Flexibly
Multiple imputation by chained equations fits a sequence of conditional models for incomplete variables.
Continuous variables can use one model.
Binary variables another.
Categorical variables another.
The flexibility is useful, but the conditional models still need to be scientifically coherent.
23. Maximum Likelihood Can Use Incomplete Cases Without Filling Every Cell Explicitly
Likelihood-based models can integrate over missing values under assumptions such as MAR.
Longitudinal mixed models often use all available outcome data without requiring complete trajectories.
This is one reason “the software accepted incomplete rows” does not mean missingness was ignored; the likelihood may be handling it under a model.
24. Full-Information Maximum Likelihood Is Common in Latent-Variable Models
Structural equation and latent-variable frameworks often use full-information maximum likelihood to estimate model parameters from all available observed data under missingness assumptions.
The approach avoids deterministic filling.
Its validity still depends on the model and missingness assumptions.
25. Inverse-Probability Weighting Reweights the Observed Cases
Suppose the probability of remaining observed can be estimated from baseline variables.
Participants who resemble people likely to drop out receive greater weight.
The weighted observed sample aims to represent the original target under assumptions.
This is inverse-probability-of-observation or censoring weighting.
26. Weighting Can Become Unstable When Observation Probabilities Are Tiny
If one participant has only a 1% estimated chance of remaining observed, their inverse probability weight can be enormous.
A few observations can dominate the analysis.
Weight diagnostics, stabilisation and positivity assumptions therefore matter.
27. Doubly Robust Methods Combine Outcome and Missingness Models
Some estimators combine a model for the outcome with a model for observation or censoring.
Under appropriate conditions they remain consistent if one of the two nuisance models is correctly specified.
“Doubly robust” does not mean robust to arbitrary violations.
The identification assumptions still matter.
28. Missing Indicators Can Be Useful Descriptively but Dangerous as a Universal Fix
A common shortcut fills missing covariates with a constant and adds a binary indicator for missingness.
This can be useful for certain prediction problems.
For causal or explanatory regression, it can bias coefficients under many realistic conditions.
A method that predicts well is not automatically a valid inferential method.
29. Structural Missingness Is Different From Accidental Missingness
A pregnancy question is not applicable to some respondents.
A follow-up question appears only if a previous answer is yes.
These values are structurally undefined rather than merely unobserved.
Encoding “not applicable” as ordinary missingness can mix different states and distort analysis.
30. Missing by Design Can Be Statistically Efficient
Some studies intentionally measure expensive variables only in a subsample.
Two-phase sampling, matrix sampling and planned missing designs can reduce cost while preserving valid inference if the missing-by-design mechanism is known and incorporated.
Not all missingness is failure.
31. Survey Nonresponse Is a Missing-Data and Sampling Problem Together
If selected people refuse to respond, the survey sample becomes incomplete.
Response propensity can depend on age, income, political interest or other variables.
Weighting and imputation often need to work together with the original sample design.
Sampling and missingness are separate stages of selection that can compound.
32. Attrition in Trials Can Break the Protection of Randomisation
Randomisation balances baseline causes at assignment.
If post-randomisation follow-up differs by treatment and outcome prognosis, the observed outcome groups can become selectively different.
Cochrane therefore treats missing outcome data as a distinct risk-of-bias domain.
Randomisation is powerful, but it cannot force missing outcomes to reappear.
33. Treatment Discontinuation and Outcome Missingness Are Not the Same
A participant can stop treatment but still provide outcome data.
Stopping the intervention should not automatically mean stopping follow-up.
Keeping outcome collection separate from treatment adherence preserves information for intention-to-treat estimands.
34. Death Can Be a Competing Event, Not an Ordinary Missing Value
If quality of life is undefined after death, the outcome is not simply “missing because the questionnaire was not returned”.
Death changes the state in which the outcome could exist.
Composite outcomes, survivor-average estimands or other strategies may be needed depending on the scientific question.
Missingness must respect the ontology of the outcome.
35. Censoring in Time-to-Event Analysis Has Its Own Assumptions
A participant may leave a study before the event occurs.
Survival analysis can accommodate censoring under assumptions about how censoring relates to future event risk.
Informative censoring can bias survival estimates if not handled appropriately.
36. Missingness Can Be Informative Even When the Missing Variable Is Not the Outcome
Suppose income is missing mainly among very wealthy participants.
Income is a confounder of exposure and outcome.
Incomplete confounder measurement can create residual confounding even if the outcome is fully observed.
Missing covariates can therefore affect causal identification.
37. MNAR Models Need Unverifiable Parameters
Because MNAR concerns unseen values, the observed data alone cannot identify every feature of the missingness mechanism.
Analysts introduce assumptions or sensitivity parameters describing how missing values might differ from MAR predictions.
The goal is not to pretend the assumption is known.
It is to show how strong the departure must be before the conclusion changes.
38. Delta Adjustment Makes Departure From MAR Explicit
Suppose MAR-based multiple imputation predicts a missing outcome of 70.
A sensitivity analysis might subtract δ = 5 from imputed outcomes for dropouts, representing the belief that unobserved outcomes are systematically worse than MAR predicts.
Vary δ across plausible values.
The analysis reveals when the scientific conclusion flips.
39. Pattern-Mixture Models Stratify by Missingness Pattern
Pattern-mixture models describe the outcome distribution within different missingness patterns and then combine them.
Because some pattern-specific outcomes are unobserved, identifying restrictions are still needed.
The structure makes the missing-data assumptions more explicit.
40. Selection Models Factor the Joint Distribution Through Missingness
Selection models specify an outcome model and a model for the probability of observation conditional on outcomes and covariates.
They provide another route to MNAR sensitivity analysis.
Different factorisations can express the same underlying joint problem from different directions.
41. Tipping-Point Analysis Asks How Bad Missing Outcomes Must Be to Reverse the Conclusion
Rather than select one unverifiable MNAR assumption, analysts can sweep across a range.
At what assumed missing-outcome disadvantage does significance disappear?
At what point does the treatment effect cross a clinically meaningful threshold?
A tipping point turns unverifiable assumptions into an interpretable robustness question.
42. Sensitivity Analysis Is Essential Because Missingness Assumptions Are Partly Untestable
One MAR analysis can look definitive.
If a modest MNAR departure reverses the conclusion, the evidence is fragile.
If the result survives a wide range of plausible departures, confidence increases.
Robustness should be measured against the assumptions we cannot verify directly.
43. Imputation Does Not Create New Information
Multiple imputation uses relationships in observed data to propagate uncertainty into plausible missing values.
It does not magically recover the actual unseen outcomes.
If no observed variable predicts the missing value and missingness depends strongly on the unseen outcome, the data remain weakly informative.
44. Machine Learning Can Impute Predictively and Still Fail Inferentially
A flexible model can predict missing values extremely well under cross-validation among observed cases.
If missing cases come from a systematically different region, prediction can fail under dataset shift.
Good predictive imputation does not remove MNAR identification problems.
45. Treating “Missing” as a Category Can Be Legitimate for Prediction
In deployed prediction systems, the fact that a value is missing can itself carry predictive information.
For example, a laboratory test may be ordered only when clinicians suspect disease.
Missingness can therefore predict outcomes.
But if clinical ordering patterns change, that predictive signal can disappear.
Missingness features can encode workflow rather than biology.
46. Missingness Can Leak Future Information
A machine-learning dataset includes whether a test result is missing.
The test was ordered only after clinicians saw later symptoms.
The missingness indicator now contains future information unavailable at prediction time.
Missing-data handling must respect the temporal boundary of the real task.
47. Education Missingness Often Contains Student State
Homework is missing because the student forgot.
Or because the task was too difficult.
Or because family circumstances interrupted study.
Or because the student has disengaged.
Treating missing work as a zero score collapses distinct states into one number.
Diagnostically, missingness can be evidence that needs interpretation rather than punishment.
48. The Hostile Test: Dropouts Are the Worst Responders
A treatment study reports excellent outcomes among completers.
Thirty percent of the treatment group dropped out because symptoms worsened.
Complete-case analysis silently removes many poor outcomes.
The observed treatment effect can be severely optimistic.
49. The Second Hostile Test: Mean Imputation Creates Artificial Certainty
Twenty missing scores are replaced with 70, the observed mean.
Those twenty values now have no variance at all.
The dataset looks more stable because uncertainty has been erased by construction.
50. The Third Hostile Test: Multiple Imputation Under an Implausible MAR Model
A sophisticated imputation routine is run with 100 imputations.
Dropout depends strongly on unrecorded symptom severity.
No sensitivity analysis is performed.
More imputations reduce Monte Carlo error.
They do not make an implausible missingness assumption true.
51. The Fourth Hostile Test: Treatment Stop Equals Outcome Stop
Participants who discontinue medication are no longer followed.
The trial therefore loses exactly the outcomes needed to estimate treatment-policy effects after discontinuation.
The operational decision created the inferential missingness.
52. The Fifth Hostile Test: “Only 3% Missing” but All From One Subgroup
Only 3% of scores are missing.
Nearly all belong to the lowest-performing rural schools.
The overall percentage is small.
The representational damage is concentrated exactly where the policy question matters.
53. Primary School: Missing Data Begins as “Why Is This Box Empty?”
A class measures plant heights.
One plant’s label falls off.
Another plant dies before measurement.
Both create blank cells.
The reasons are different.
A missing value has a story. Find the story before deciding what to do with the blank.
54. Secondary School: Missingness Becomes Selection
Students can compare a full class average with an average calculated only among students who attended the hardest test.
If absence is related to ability, the observed group is selected.
The lesson is simple:
who disappears from the dataset can change what the dataset says.
55. JC and University: Missing Data Becomes an Identification and Sensitivity Problem
At higher levels, learners should reconstruct:
- estimand;
- missingness indicator;
- missingness pattern;
- MCAR/MAR/MNAR assumptions;
- observed predictors of value and missingness;
- complete-case assumptions;
- multiple imputation;
- likelihood methods;
- inverse-probability weighting;
- MNAR sensitivity parameters;
- tipping points;
- structural missingness;
- censoring;
- uncertainty propagation.
The blank cell becomes a node in the causal and inferential system.
56. Where Missing Data Fits in the eduKateSG “How Works” Landscape
- How Observation Works — how world states become records and how missingness enters observation.
- How Sampling Works — who enters the dataset before later attrition.
- How Research Bias Works — missing-outcome and selection bias.
- How Statistical Inference Works — the wider estimand, model and uncertainty system.
- How Confidence Intervals Work — uncertainty that missing-data methods must propagate.
- How Randomisation Works — assignment protection that attrition can later weaken.
- How Research Variables Work — outcome, covariate, mediator and missingness roles.
Missing Data owns one precise canonical job: explain how unobserved values arise, what assumptions connect them to observed information, which methods can recover the target estimand under those assumptions, and how sensitive the conclusion is when the assumptions cannot be verified.
57. What This Article Does Not Claim
- Missing data are not automatically harmless when the percentage missing is small.
- MAR does not mean values are missing randomly in ordinary language.
- MAR cannot generally be proven from observed data alone.
- MCAR/MAR/MNAR labels do not by themselves determine the best method in every multivariable problem.
- Complete-case analysis is not assumption-free.
- Mean or deterministic single imputation usually understates uncertainty.
- Multiple imputation does not recover the actual missing values or make MNAR assumptions disappear.
- More imputations reduce simulation error, not identification uncertainty.
- Machine-learning prediction of missing values does not automatically produce valid causal inference.
- Prevention and continued outcome collection are often more valuable than sophisticated repair after data are lost.
58. A Compact Missing-Data Audit
- What target estimand is being estimated?
- Which variables are incomplete?
- How much data are missing overall?
- How is missingness distributed across treatment groups, sites and subgroups?
- Is the pattern monotone or intermittent?
- What operational reasons caused missingness?
- Which observed variables predict missingness?
- Which observed variables predict the missing values?
- Is MCAR plausible?
- What MAR assumption is being made?
- What MNAR mechanisms are plausible?
- Was missingness prevented where possible?
- Were outcomes collected after treatment discontinuation?
- Is complete-case analysis used, and under what assumption?
- Is multiple imputation appropriate?
- Does the imputation model include outcome, exposure and useful auxiliary variables?
- Does it preserve interactions, nonlinearities and clustering?
- Are Rubin’s rules or another valid combination method used?
- Could likelihood methods use incomplete observations directly?
- Would weighting better represent the observation process?
- Are weights stable and positivity plausible?
- What sensitivity analysis addresses MNAR departures?
- What tipping point changes the conclusion?
- Are structurally inapplicable values separated from accidental missingness?
- Is censoring informative?
- Does the final uncertainty include missing-data uncertainty?
59. Frequently Asked Questions
What is missing data?
Missing data are values relevant to an analysis that were not observed or recorded for some study units. The statistical problem depends on why the values are missing and what information remains available.
What is MCAR?
Missing completely at random means the probability a value is missing does not depend on the observed or unobserved values relevant to the analysis under the specified data structure.
What is MAR?
Missing at random means that, after conditioning on the relevant observed information included in the missing-data model, missingness does not additionally depend on the unseen value itself.
What is MNAR?
Missing not at random means the probability of missingness still depends on unobserved values or unobserved factors after accounting for the observed information. Such mechanisms require untestable assumptions and sensitivity analysis.
Is multiple imputation always best?
No. Multiple imputation is powerful under appropriate assumptions, but likelihood methods, weighting, planned-design estimators or other approaches may be better for particular estimands and data structures. The method must match the missingness process and scientific target.
60. Authoritative Research Corridor
- SPIRIT 2025 Explanation and Elaboration — Prevention and Handling of Missing Outcome and Covariate Data
- Cochrane Handbook Chapter 8 — Bias Due to Missing Outcome Data
- Assumptions and Analysis Planning in Studies With Missing Data in Multiple Variables: Moving Beyond MCAR/MAR/MNAR
- Missing Data: A Statistical Framework for Practice
- Missing Data Methods in Longitudinal Studies: A Review
Final Thought: The Blank Is Part of the Evidence
Researchers are trained to look at numbers.
Missing-data analysis teaches us to look at absence.
Who stopped answering?
Which sensor failed?
Which student disappeared from the final test?
Which patient could no longer attend?
What does that disappearance tell us about the unseen value?
The best method is not the one that makes every spreadsheet cell look complete.
It is the one that preserves uncertainty honestly while using all defensible information the study still has.
Missing data are not nothing. They are evidence about the limits of observation—and sometimes evidence about the very process we are trying to understand.