HEW-NODE-0136 · How Education Works · impact evaluation, causal inference, counterfactuals, randomised evaluations, quasi-experiments, difference-in-differences, regression discontinuity, instrumental variables, matching, selection bias, implementation evidence, spillovers, heterogeneity, external validity, cost evidence and policy learning
A programme begins. Test scores rise. Attendance improves. Parents are happier. The minister says the reform worked.
Perhaps it did.
But perhaps the cohort was stronger, the economy improved, another policy arrived at the same time, schools volunteered because they were already better organised, the weakest students left the measured sample, teachers became more experienced, or the outcome had been rising for three years before the programme began.
The difficult question is not whether the outcome changed. It is whether the outcome changed because of the intervention.
Impact evaluation is the discipline of constructing a credible answer to the question education policy most often skips: what would have happened to these learners, schools or systems if the reform had not occurred?
This node sits beside the How Education Works hub, Educational Research, Education Sector Analysis & System Diagnosis, Education Policy Pilots & Scaling, Education Statistics Quality Assurance & Data Validation, Education Cost-Effectiveness & Benefit-Cost Analysis, Education Benchmarking & Policy Transfer and Education Research Governance, Ethics & Data Access.
Those pages keep their jobs. Educational Research owns the wider landscape of research questions, methods and evidence. Sector Analysis owns diagnosis of how the system is functioning. Policy Pilots owns bounded implementation before scale. Statistics QA owns the quality of administrative and survey data. Cost-Effectiveness owns comparison of costs and outcomes. This node owns the causal claim itself: how education systems estimate whether a defined intervention changed an outcome relative to a credible counterfactual, how they test competing explanations, and how they decide how far the finding can travel.
The 60-Second Read
- Before-and-after change is not automatically impact.
- A counterfactual is an estimate of what would have happened without the intervention.
- The central causal problem is that the same learner cannot be observed both treated and untreated at the same time.
- Random assignment can create comparable groups before treatment, but only when it is implemented properly and analysed as designed.
- Quasi-experimental methods can estimate causal effects when randomisation is unavailable and design assumptions are credible.
- Difference-in-differences relies on a defensible comparison of trends, not merely two averages.
- Regression discontinuity can exploit assignment thresholds when units close to the threshold are genuinely comparable.
- Instrumental-variable designs require an instrument that changes treatment without affecting the outcome through another route.
- Matching reduces observed differences; it cannot repair unobserved confounding by magic.
- Selection into a programme is often the main source of false causal stories.
- Attrition can recreate selection after a good evaluation begins.
- Spillovers can contaminate comparison groups and also represent real system effects worth measuring.
- Implementation failure and theory failure are different diagnoses.
- An average effect can hide winners, losers and null effects across groups or settings.
- External validity asks whether an effect estimated somewhere else will survive a new context, population, delivery system or scale.
- Statistical significance does not tell policymakers whether an effect is educationally important.
- A small effect at low cost can outperform a larger effect with huge cost or implementation burden.
- Null findings can prevent waste and deserve protection from publication or political bias.
- Impact evaluation should be planned before rollout when possible, because design options disappear after implementation begins.
- The goal is not to make every policy an experiment. It is to make important causal claims earn the confidence attached to them.
One-Sentence Definition
Education impact evaluation is the systematic estimation of the causal effect of an intervention by comparing observed outcomes with a credible estimate of the outcomes that would have occurred in the absence of that intervention.
The First Distinction: Change Is Not Impact
Suppose reading scores rise from 470 to 485 after a new literacy programme. That is a change of fifteen points. It is not yet an impact of fifteen points.
If similar schools without the programme also rose from 468 to 481 because curriculum materials improved nationally, then part of the observed gain would likely have happened anyway. Impact is the difference between the observed treated outcome and the best credible estimate of the untreated outcome.
The counterfactual is invisible, so evaluation design creates evidence about it.
The Second Distinction: Correlation Is Not Causal Attribution
Students who attend tutoring may score higher than students who do not. That pattern can arise because tutoring helps, because more motivated families purchase tutoring, because struggling students seek tutoring, or because tutoring access is concentrated in particular schools.
A correlation can be important evidence while remaining insufficient for the causal question.
The Third Distinction: Programme Performance Is Not Programme Impact
A programme can meet every implementation target — train 10,000 teachers, distribute every device, hold every workshop — and still produce little change in learning. Another programme can miss some process targets while creating meaningful outcomes.
Performance monitoring asks whether planned activities and outputs happened. Impact evaluation asks what those activities caused.
Current World Bank Definition: The Counterfactual Is the Core
The World Bank’s current compendium of education-focused impact evaluations defines impact evaluation around rigorous experimental or quasi-experimental methods that compare a group receiving an intervention with a credible comparison group. Its education evidence programme includes studies published from 2015 to 2024 and focuses on learning outcomes.
The World Bank’s DIME Education programme extends the same logic into policy partnerships: rigorous causal evidence is combined with diagnostics, systematic reviews, administrative data and implementation support so that governments can learn not only whether something worked, but under what conditions.
The Fundamental Problem of Causal Inference
For one student, imagine two potential outcomes:
- the score if the student receives the programme;
- the score if the same student, at the same moment, does not receive it.
Only one can be observed. Once the programme is assigned, the alternative outcome becomes counterfactual. Evaluation therefore uses other students, schools, thresholds, timing or naturally occurring variation to estimate the missing comparison.
Begin With the Causal Question
“Did the programme work?” is too broad. A stronger question specifies:
- intervention;
- population;
- comparison condition;
- outcome;
- time horizon;
- implementation setting;
- causal estimand — whose effect, exactly, is being estimated.
For example: What was the effect of two years of the new lower-secondary mathematics coaching programme on Grade 9 mathematics achievement among schools offered the programme in the first rollout wave, relative to similar schools scheduled for later rollout?
Draw the Theory of Change Before Selecting the Method
An evaluation should know how the programme is supposed to create the outcome.
resources → programme activity → changed teacher or learner behaviour → intermediate outcome → final outcome
If the programme is teacher coaching, the evaluator might expect coaching sessions to improve instructional practice, which increases productive student practice, which then improves learning. Measuring only final scores leaves the mechanism invisible.
A Logic Model Prevents the Wrong Null
If test scores do not improve, three stories are possible:
- the programme was not delivered;
- the programme was delivered but did not change the intended behaviour;
- the behaviour changed but did not affect the measured outcome.
These imply different decisions. Implementation evidence is therefore not optional decoration around an impact estimate.
Randomisation Solves Selection Before Treatment
When eligible schools or students are randomly assigned to treatment and comparison groups, treatment status is independent of pre-existing characteristics in expectation. That allows average outcome differences after treatment to estimate the causal effect under the design.
The World Bank’s DIME methodological guidance describes experimental variation as a way to avoid confounding because treatment assignment is generated independently of the factors that otherwise bias comparisons.
Randomisation Is a Design, Not a Magic Word
A study can call itself randomised and still fail operationally. Treatment may be reassigned, schools may cross over, participants may refuse, baseline records may be wrong, or implementers may influence assignment.
Preserve the randomisation record, compare baseline balance, document deviations and analyse according to the pre-specified assignment logic.
Randomise at the Level Where Spillovers Are Manageable
If teachers in the same school share lesson materials, randomising individual teachers may contaminate the comparison group. Cluster randomisation at school level may better preserve the treatment contrast, though it requires more schools and changes the statistical analysis.
Power Is a Design Constraint
An evaluation with too few schools can miss educationally meaningful effects because estimates are imprecise. Adding thousands of students inside a small number of highly similar clusters does not always solve the problem because treatment varies at the school level.
Power calculations should match assignment level, expected effect size, outcome variability, clustering and attrition.
Intent-to-Treat Protects the Assignment
If some assigned schools do not participate fully, comparing only compliant schools can reintroduce selection. Intent-to-treat analysis compares groups according to original assignment, estimating the effect of being offered or assigned the programme under real implementation conditions.
Other estimands can examine treatment received, but they require additional assumptions.
Waiting Lists Can Create Ethical Evaluation Designs
When a programme cannot reach every eligible school immediately, random order of phased rollout can create a comparison group without permanently denying access. Early and later waves can be compared during the rollout period.
Quasi-Experiments Use Structure Already in the World
Randomisation is not always possible. Governments rarely randomise national laws, school-entry ages or historical funding reforms. Quasi-experimental designs use thresholds, timing, geography, policy rules or exogenous shocks to create comparisons that approximate experimental logic.
The word quasi does not mean weak. A well-designed natural experiment can be more credible than a badly implemented randomised trial.
Difference-in-Differences Uses Change Relative to Change
Suppose one region introduces a policy and a similar region does not. If the treated region improves by eight points and the comparison region improves by three, a simple difference-in-differences estimate attributes the extra five-point change to treatment under the design assumptions.
The key assumption is not that the regions had identical baseline levels. It is that, without treatment, their outcome trends would have evolved comparably.
Pre-Trends Are a Diagnostic, Not a Proof
Similar trends before treatment increase confidence in a difference-in-differences design, but they cannot prove that future untreated trends would have remained parallel. Policy timing, anticipation and concurrent shocks still need investigation.
Staggered Adoption Needs Modern Methods
When districts adopt a policy at different times, old two-way fixed-effect models can mix treated groups into comparison groups in ways that distort estimates when effects change over time. Modern event-study and cohort-specific approaches can better respect treatment timing.
The lesson for policy users is simple: “difference-in-differences” describes a family of designs, not one automatic spreadsheet formula.
Regression Discontinuity Uses a Threshold
If students above a score threshold receive a scholarship while students just below do not, outcomes near the threshold can sometimes be compared because the groups may be similar except for treatment assignment.
The estimate is local: it tells us about units near the cutoff. It should not automatically be generalised to students far away from it.
Check for Manipulation Around the Threshold
If schools can move students across the eligibility cutoff strategically, the design weakens. Inspect score density, assignment procedures and whether actors knew the rule before measurement.
Instrumental Variables Need an Exclusion Story
An instrumental variable changes treatment participation but should affect the outcome only through that treatment. For example, random programme assignment might be used as an instrument for actual participation when not everyone complies.
The mathematical estimate is only as credible as the substantive argument that the instrument has no other route to the outcome.
Matching Controls What You Can Observe
Propensity-score or covariate matching can construct comparison groups with similar observed characteristics. That can improve comparability, especially when treatment assignment depends on recorded variables.
Matching cannot guarantee balance on motivation, leadership quality or other unobserved factors. It should not be described as if it creates randomisation after the fact.
Interrupted Time Series Uses a Long Run of Data
When a national policy begins at a known time, repeated outcome measurements before and after implementation can test whether the level or trend changes unusually at the intervention point.
Concurrent national shocks remain a threat because there may be no untreated comparison. Controlled time-series designs strengthen the claim when another unaffected series is available.
Synthetic Controls Construct a Comparison From Several Places
If one region implements a large reform and no single comparison region is suitable, analysts can sometimes combine weighted information from several unaffected regions to reproduce the treated region’s pre-policy trajectory.
The resulting synthetic comparison is transparent about the donor pool and weights, but credibility still depends on whether untreated units can reproduce the relevant counterfactual.
Selection Bias Is the Default Enemy
Education programmes often attract participants systematically different from non-participants. Strong principals volunteer for pilots. Wealthier families select schools. Motivated teachers attend professional development. Struggling students join remediation.
Selection can bias estimates upward or downward. The direction should never be assumed.
Regression to the Mean Can Look Like Repair
Students are often selected for intervention after an unusually low score. Even without treatment, some would score closer to their usual level next time. A before-after improvement among low scorers can therefore overstate impact.
Maturation Can Look Like Programme Effect
Young learners improve with age and ordinary schooling. An intervention delivered during a year of rapid development needs a comparison group because natural growth can be large.
History Can Confound the Evaluation
A new curriculum, examination reform, economic shock or teacher-pay change can coincide with the programme. Evaluation should map concurrent events and ask which groups were exposed.
Measurement Change Can Manufacture Impact
If the post-test is easier, more aligned to the intervention or scored differently, observed gains can reflect measurement rather than learning. Stable outcome measurement is part of causal identification.
Teaching to the Measure Can Be a Real Effect and a Validity Problem
A programme may improve performance on the exact assessment used for evaluation while producing little transfer to broader capability. Evaluate near-transfer and farther-transfer outcomes where the policy claim extends beyond the trained tasks.
Attrition Can Break the Original Comparison
If weaker students disappear disproportionately from one group, the remaining sample is no longer comparable. Report attrition by group, investigate reasons and use sensitivity analysis rather than silently analysing whoever remains.
Missing Data Are Part of the Result
A programme that makes outcome data less likely to be collected can bias results. Missingness should be described, modelled where justified and treated as an implementation signal.
Spillovers Can Make the Treatment Look Smaller
Teachers in comparison schools may receive materials from colleagues in treatment schools. Students may share resources. District leaders may copy practices. The comparison group becomes partially treated.
That weakens the direct contrast but can reveal an important system mechanism: the intervention diffuses.
Spillovers Can Also Be the Outcome We Care About
Peer effects, teacher collaboration and community information can create benefits beyond direct recipients. Evaluation should sometimes estimate total system effect rather than treating every spillover as contamination.
Displacement Can Hide Costs Elsewhere
If a scholarship programme improves outcomes for recipients partly by drawing strong students away from other institutions, recipient effects do not equal net system effects. A school-level intervention can reallocate teachers or leadership attention away from other grades.
Impact evaluation should define whether it estimates participant, provider or system impact.
Implementation Evidence Explains the Path
Track dosage, participation, quality, timing, adaptation and reach. An intervention delivered at half the planned intensity is a different intervention from the protocol described in the policy document.
Fidelity Is Not Blind Uniformity
Some programmes require adaptation to local conditions. Separate core components that carry the hypothesised mechanism from adaptable components that can vary safely.
Mechanism Measures Make Evaluation More Useful
If coaching improves student achievement, measure teacher practice. If a transport subsidy improves attendance, measure journey constraints. If a grant improves completion, measure whether liquidity, study time or course choice changed.
Mechanism evidence turns “worked” into a more transferable statement.
Average Effects Can Hide Heterogeneity
A programme can help novice teachers and do little for experienced teachers, benefit low-performing schools and burden high-performing schools, or work in urban areas while failing where travel is difficult.
Subgroup analysis should follow theory and be pre-specified where possible. Searching hundreds of subgroups after the fact can generate misleading “discoveries.”
Distributional Effects Matter Even When the Mean Is Positive
An intervention can raise average performance while widening inequality. Report effects across relevant groups and, where possible, across the outcome distribution.
Statistical Significance Is Not Educational Importance
A tiny effect can become statistically significant in a huge sample. A meaningful effect can be statistically uncertain in a small study. Policymakers need effect size, uncertainty interval, baseline context and practical consequence.
Confidence Intervals Are Part of the Finding
An estimate of +0.12 standard deviations with a narrow interval tells a different story from +0.12 with an interval spanning substantial benefit and harm. Report the range the data can reasonably support.
Null Does Not Mean Zero
A statistically non-significant result can mean the effect is small, the study is underpowered, the implementation was weak or the estimate is imprecise. Interpret null findings using confidence intervals and implementation evidence rather than writing “no effect” automatically.
Negative Findings Are System Assets
If a costly programme produces little effect, publishing that result can prevent repetition elsewhere. Political systems often reward positive announcements more than disciplined learning; evaluation governance should protect inconvenient findings.
Pre-Analysis Plans Reduce Outcome Shopping
When hundreds of outcomes and model specifications are available, analysts can find apparently positive results by chance. Pre-specifying primary outcomes, models and subgroup tests makes the distinction between confirmatory and exploratory analysis clearer.
Multiple Testing Needs Control
If twenty independent outcomes are tested at the five-per-cent significance threshold, false positives become likely. Group outcomes into families, control error rates where appropriate and avoid presenting one isolated significant coefficient as the whole story.
Specification Curves Can Expose Fragility
When several reasonable modelling choices exist, show whether the conclusion survives those choices. An effect that appears only under one narrow specification deserves weaker confidence.
Replication Is More Than Re-Running the Same Code
Computational replication checks whether the reported estimate can be reproduced from the same data and code. Conceptual replication asks whether the effect appears in another sample, implementation team, year or setting. Both matter for policy.
External Validity Is a Mechanism Question
A tutoring programme may work in a trial with highly trained tutors, small groups and strong monitoring. Scaling through ordinary staffing can change all three conditions. The question is not simply “Does the result generalise?” but “Which causal ingredients will survive the new delivery system?”
Population Transport Requires More Than Demographic Similarity
Two student populations can have similar age and income while schools differ in teacher capacity, curriculum, attendance, language or institutional incentives. External validity depends on moderators that interact with the intervention mechanism.
Scale Can Change the Treatment
A small pilot can recruit unusually strong staff, receive intensive support and operate without changing market wages. At national scale, tutor supply, textbook prices, teacher labour markets and administrative capacity can respond.
Scale is not merely “more units receiving the same thing.”
General Equilibrium Effects Belong in System Evaluation
A scholarship programme can increase demand for tertiary places and raise fees. Teacher incentives can shift teachers between schools. Expanding early-childhood provision can change labour supply of caregivers and parents.
Large policies can change the environment in which the original causal estimate was produced.
Cost Belongs Next to Effect
An effect of +0.15 standard deviations says nothing about whether the programme costs $20 or $2,000 per student. Connect impact estimates to Education Cost-Effectiveness & Benefit-Cost Analysis before comparing alternatives.
Implementation Burden Is a Cost
A programme can have a reasonable financial cost and an unsustainable administrative burden. Record teacher time, reporting load, training requirements and management capacity.
Ethics Begins Before Randomisation
Not every policy question should be randomised. Denying an established essential service, exposing learners to known harm or concealing material risks can be unethical. Equipoise, consent, gatekeeper authority, safeguarding and data governance matter.
See Education Research Governance, Ethics & Data Access for the research-authority layer.
Fairness Can Change the Evaluation Design
When places are scarce, randomisation among equally eligible applicants can sometimes allocate access fairly while also enabling evaluation. When need differs substantially, weighted priority may be more ethical than equal lottery.
Administrative Data Can Make Evaluation Cheaper
Student records, attendance, assessment and teacher data can reduce data-collection cost and support long follow-up. They also create privacy, linkage, missingness and measurement-quality issues. Data originally collected for administration do not become research-ready automatically.
Outcome Quality Matters More Than Outcome Convenience
If the programme aims to improve reasoning, a convenient administrative multiple-choice score may capture only part of the goal. Use outcomes aligned to the causal claim and identify what each measure omits.
Long-Term Follow-Up Can Reverse the Story
Some interventions create immediate gains that fade. Others produce modest short-run effects and larger long-run impacts through persistence, progression or later choices. Evaluation timing should match the theory of change.
Fade-Out Does Not Automatically Mean No Value
A temporary gain can affect grade progression, confidence, course access or another intermediate decision even if test-score differences later narrow. Track pathways, not only one repeated score.
Mediation Analysis Needs Stronger Assumptions
It is tempting to say teacher practice “mediated” the effect because practice changed and learning changed. Causal mediation requires additional assumptions about confounding of the mediator-outcome relationship. Mechanism claims should be proportionate to the design.
Qualitative Evidence Can Explain Causal Process
Interviews and observation can reveal why schools used the intervention differently, why teachers resisted one component or why families interpreted the programme unexpectedly. Qualitative evidence may not identify the average causal effect, but it can make the effect interpretable.
Mixed Methods Are Strongest When Each Method Has a Job
Do not add interviews merely to make a study “mixed.” Use quantitative identification to estimate effect, implementation measures to verify delivery, and qualitative work to inspect mechanism, adaptation and experience.
Monitoring Data Can Trigger Adaptive Evaluation
If implementation data reveal that half the sites never received the intervention, the evaluation should preserve the original causal estimate while adding analyses that clarify treatment exposure. Do not silently redefine treatment after seeing outcomes.
Evaluation Governance Should Protect Independence
The programme team understandably wants success. Evaluators need enough independence to report weak, null or adverse findings. Conflicts, sponsor roles, data access and publication rights should be explicit.
Decision Rules Should Be Set Before the Result
If a ministry says it will scale any programme with a positive estimate, it may scale effects too small to matter. If it requires statistical significance alone, it can reject useful but imprecisely estimated interventions.
Predefine a decision framework using effect size, uncertainty, cost, implementation quality, equity, strategic fit and downside risk.
Evidence Thresholds Can Differ by Decision
A reversible low-cost pilot can proceed with weaker evidence than a national high-cost mandate or irreversible infrastructure decision. Evidence requirements should rise with consequence.
Use Evidence Synthesis Before Starting Another Trial
A new evaluation adds little if twenty high-quality studies already answer the same question in similar settings. Systematic reviews and meta-analyses can show where uncertainty actually remains.
But Meta-Analysis Does Not Remove Context
Average effects across studies can hide variation in dosage, age, subject, delivery model and baseline conditions. Ask what predicts heterogeneity before importing the pooled mean.
Policy Learning Needs an Evidence Ledger
- programme version;
- target population;
- theory of change;
- evaluation design;
- assignment rule;
- primary outcomes;
- implementation measures;
- effect estimates;
- uncertainty;
- subgroup results;
- spillovers;
- cost;
- implementation burden;
- external-validity conditions;
- decision taken;
- later replication or scale evidence.
Case Study: The Programme That Improved Before It Began
Invented example: districts adopting a new attendance programme show a six-point improvement after launch. Historical data reveal those districts had already been improving faster for two years because they had stronger leadership teams.
A naïve before-after comparison credits the programme. A trend-aware design shows little additional change beyond the pre-existing trajectory. The policy team redesigns the intervention rather than scaling a story.
Case Study: The Threshold Scholarship
Invented example: students scoring 70 or above receive a fee waiver. Students at 69 and 70 are similar in observed background but face different eligibility. A regression-discontinuity design estimates the scholarship effect near the threshold.
The finding is useful for borderline-eligible students. It does not automatically establish the effect for students scoring 40 or 95.
Case Study: The Trial That Became a Different Programme at Scale
Invented example: a tutoring trial uses university graduates, groups of three, weekly supervision and detailed curriculum alignment. The national rollout uses temporary staff, groups of eight and minimal supervision because supply cannot keep pace.
The scale result differs. The original evaluation was not necessarily wrong. The treatment changed.
Case Study: The “No Effect” Programme That Was Never Delivered
Invented example: a teacher professional-learning programme shows no test-score effect. Implementation data reveal only 38 per cent of teachers received the planned coaching cycle and rural schools received half the intended visits.
The policy conclusion becomes “implementation model failed under these conditions,” not “coaching can never work.”
Case Study: The Positive Average That Hid Exclusion
Invented example: a digital remediation platform raises average scores. Device logs show the weakest students use it least because home connectivity is poor. High-use students drive the average gain.
The system retains the intervention but changes access design and evaluates effects by baseline achievement and connectivity status.
Failure Mode 1: Before-and-After = Impact
Repair: construct a credible counterfactual using design rather than optimism.
Failure Mode 2: Volunteer Schools Compared With Everyone Else
Repair: address selection through randomisation, assignment rules or a stronger quasi-experimental design.
Failure Mode 3: Matching Treated as Randomisation
Repair: state that matching balances observed variables and test sensitivity to unobserved confounding.
Failure Mode 4: Attrition Ignored
Repair: report missing outcomes by group and examine whether dropout recreates selection.
Failure Mode 5: Implementation Not Measured
Repair: track dosage, quality, reach and adaptations so null effects can be diagnosed.
Failure Mode 6: One Significant Outcome Becomes the Headline
Repair: pre-specify primary outcomes, control multiple testing and report the full outcome family.
Failure Mode 7: Average Effect = Effect for Everyone
Repair: examine theoretically justified heterogeneity and distributional effects.
Failure Mode 8: Trial Effect = Scale Effect
Repair: compare staffing, dosage, delivery capability, market responses and institutional conditions at scale.
Failure Mode 9: Statistical Significance = Policy Worth
Repair: integrate magnitude, uncertainty, cost, equity and implementation burden.
Failure Mode 10: Inconvenient Null Findings Disappear
Repair: register evaluations, protect publication rights and store findings in the policy evidence base.
The Education Impact-Evaluation Operating Chain
- Define the policy decision.
- Define the causal question.
- Specify population, intervention, comparison, outcome and time horizon.
- Map the theory of change.
- Identify core implementation mechanisms.
- Review existing evidence.
- Decide whether new causal evaluation is needed.
- Select randomised or quasi-experimental design.
- Pre-specify primary outcomes and estimands.
- Conduct power analysis.
- Set assignment or eligibility rules.
- Register the design where appropriate.
- Collect baseline data.
- Verify assignment integrity.
- Monitor treatment delivery.
- Track contamination and spillovers.
- Track attrition and missingness.
- Protect outcome measurement integrity.
- Estimate primary effects.
- Report uncertainty.
- Conduct design diagnostics.
- Test pre-specified heterogeneity.
- Analyse mechanism evidence.
- Assess adverse or unintended effects.
- Connect effect estimates to cost.
- Assess external-validity conditions.
- Compare findings with prior evidence.
- Make a decision using a pre-defined evidence framework.
- Store data, code and documentation under appropriate governance.
- Re-evaluate after adaptation or scale.
An Impact-Evaluation Dashboard
- intervention version;
- population;
- assignment unit;
- sample size;
- primary outcomes;
- baseline balance;
- treatment take-up;
- implementation dosage;
- comparison contamination;
- attrition by group;
- primary effect estimate;
- confidence interval;
- effect size in natural units;
- subgroup effects;
- spillovers;
- adverse effects;
- cost per participant;
- cost-effectiveness;
- external-validity conditions;
- decision taken;
- later replication or scale result.
Canonical Owner Boundaries
- Educational Research owns research methods and evidence broadly.
- Education Sector Analysis & System Diagnosis owns descriptive and diagnostic analysis of the education system.
- Education Policy Pilots & Scaling owns bounded implementation and scale-up mechanics.
- Education Statistics Quality Assurance & Data Validation owns the quality and validation of education statistics.
- Education Cost-Effectiveness & Benefit-Cost Analysis owns comparison of costs and benefits after outcome evidence exists.
This node owns the causal bridge: the design, assumptions and evidence required to move from “the outcome changed after the reform” to “the reform changed the outcome relative to a credible counterfactual.”
The Return Path
Return to the graph that rises after a reform begins.
It is tempting to stop there. The line moved in the desired direction. The intervention has a name. The timing looks persuasive.
But education systems make decisions at scale. A false causal story can consume years of teacher time, millions in public money and an entire cohort’s opportunity to receive something better.
Impact evaluation exists because the most expensive question is often not “Did something change?”
It is “Would it have changed anyway?”
A credible counterfactual is one of the quietest pieces of infrastructure in good education policy: it stops coincidence from becoming strategy.
Return to the How Education Works hub.