HEW-NODE-0157 · How Education Works · evidence standards, evidence clearinghouses, systematic review, causal evidence, replication, evidence synthesis, research quality, evidence ratings, practice guides, implementation evidence, external validity, context transfer, policy use and research transparency
An education ministry can have too much research and still lack usable evidence.
One study reports a large positive effect. Another finds almost nothing. A vendor cites a university trial. A teacher association shares a case study. A systematic review averages results across countries. A pilot works in one district. A preprint arrives before peer review. A meta-analysis is rigorous but mostly includes interventions unlike the one being considered.
The problem is no longer access to research. The problem is deciding what the evidence can legitimately support.
An evidence clearinghouse is public decision infrastructure: it separates finding research from judging research, and judging research from deciding whether it applies here.
This node sits beside the How Education Works hub, Educational Research, Education Impact Evaluation & Causal Inference, Education Research Governance, Ethics & Data Access, Education Cost-Effectiveness & Benefit-Cost Analysis and Education Policy Implementation, Adaptive Delivery & Monitoring.
Those pages keep their jobs. Educational Research owns research design and inquiry broadly. Impact Evaluation owns causal identification. Research Governance owns ethics and data access. Cost-Effectiveness owns value-for-money comparison. Implementation owns delivery. This node owns the evidence-qualification layer: how an education system reviews studies against explicit standards, synthesises bodies of research, communicates certainty, distinguishes evidence of effect from evidence of implementation, and turns research into decision-grade guidance without erasing uncertainty or context.
The 60-Second Read
- A research paper is not automatically policy-grade evidence.
- Study design, execution and reporting determine what causal claims are justified.
- Evidence standards should be published before politically attractive findings appear.
- Clearinghouses review studies using common rules so users do not have to reinvent appraisal each time.
- One strong study can establish important evidence and still leave transfer uncertainty.
- Multiple weak studies do not automatically become strong evidence when averaged.
- Systematic reviews need explicit inclusion rules and transparent search methods.
- Meta-analysis summarises evidence but cannot repair poor underlying studies.
- Independent replication matters because developer-led studies can differ from external evaluations.
- Effect size and evidence security answer different questions.
- A large estimated effect with weak evidence may deserve more testing rather than rapid scale.
- Null findings are evidence too.
- Heterogeneity matters: average effects can hide major differences by age, subject, setting or implementation.
- External validity asks whether evidence travels to the new population and system.
- Implementation evidence asks whether the intervention can be delivered as intended.
- Cost evidence asks whether the intervention is feasible at scale.
- Evidence ratings should explain why confidence is high or low rather than merely display a label.
- Practice guides should distinguish research-backed recommendations from expert judgement.
- Evidence needs freshness rules because programmes, curricula and technology change.
- The objective is not to find research that supports a decision. It is to build a system that makes disconfirming evidence hard to ignore.
One-Sentence Definition
An education evidence clearinghouse is a structured system that applies explicit research standards to individual studies and bodies of evidence, then translates the resulting confidence, effects, limitations and implementation conditions into forms that educators and policy-makers can use.
The First Distinction: Evidence Quality Is Not Effect Size
An intervention can report a large positive effect from a weak design. Another can report a modest effect from a large, well-executed randomised trial. Effect magnitude tells us what was estimated. Evidence quality tells us how much confidence the design earns.
A clearinghouse should keep those dimensions separate.
The Second Distinction: Evidence of Effect Is Not Evidence of Transfer
A tutoring model can work in one age group, subject, staffing model and country. That does not automatically establish the same effect in another system.
Internal validity asks whether the study identified an effect credibly. External validity asks whether that evidence is relevant to the new decision context.
The Third Distinction: Evidence Clearinghouse Is Not a Search Engine
A search engine retrieves documents. A clearinghouse adds judgement rules: which studies qualify, what design standards apply, how findings are coded, how multiple studies are synthesised and what users should infer.
The difficult work is not retrieval. It is disciplined interpretation.
Apex Example: The What Works Clearinghouse
The US Institute of Education Sciences’ What Works Clearinghouse reviews education programmes, products, practices and policies using published procedures and evidence standards. Its current Standards Version 5.0 sets rules for reviewing study designs and implementation, while its products include individual study reviews, intervention reports and practice guides.
The institutional principle is more important than any one national framework: evidence review should be governed by rules that are visible before the result is known.
Standards Should Be Pre-Committed
If standards change depending on whether the result supports a preferred reform, evidence review becomes advocacy. Define acceptable designs, attrition thresholds, baseline-equivalence rules, outcome requirements, statistical adjustments and reporting standards in advance.
Pre-commitment does not make a standard timeless. It makes revisions traceable.
Different Questions Need Different Evidence Designs
Randomised trials are powerful for many causal questions. They are not the answer to every education question. Long-term policy effects, system reforms, rare harms, implementation processes and historical changes may require quasi-experimental, observational, qualitative or mixed-method evidence.
A clearinghouse should match evidence design to the claim rather than reduce research quality to one hierarchy used for every purpose.
Study Review Needs More Than the Design Label
“Randomised trial” can sound decisive, but execution matters. Was randomisation preserved? Was attrition differential? Were outcomes selected after results were known? Was analysis consistent with assignment? Were clusters handled correctly? Were baseline differences material?
Evidence standards turn those questions into reproducible review procedures.
Review Protocols Define the Question Before the Literature Is Seen
A systematic review should define population, intervention, comparison, outcomes, study designs, dates, languages and search strategy before selecting studies. Otherwise inclusion can drift toward convenient findings.
Publication Bias Distorts the Visible Literature
Positive and statistically significant findings are more likely to be published than null findings in many research environments. A clearinghouse should search beyond prominent journal articles where possible, including registered trials, reports and other eligible sources.
Developer-Led Evaluation Needs Special Attention
The organisation that created an intervention may have strong expertise and legitimate reasons to evaluate it. It may also have incentives, implementation support or researcher involvement that differ from ordinary scale-up.
The Education Endowment Foundation’s current Teaching and Learning Toolkit explicitly treats lack of independent evaluation as one factor that can lower evidence security. The point is not that developer-led evidence is invalid. It is that independence changes what we learn about generalisability.
Evidence Security and Estimated Impact Should Be Separate
The EEF Toolkit communicates both estimated impact and evidence strength. That is a valuable design principle. Users can see that an approach may look promising while the underlying evidence remains limited, or that a moderate average effect rests on a large evidence base.
Confidence should never be smuggled inside the effect estimate.
Heterogeneity Is Information, Not Noise to Hide
Two hundred studies can produce one average number and still contain major variation. Effects may differ by age, subject, dosage, implementation, prior attainment, delivery role or context.
A strong synthesis asks what explains variation and whether the average meaningfully describes the new setting.
Meta-Analysis Does Not Automatically Increase Truth
Meta-analysis combines estimates. If included studies are biased, incomparable or measure different constructs, statistical aggregation can produce a precise-looking answer to an unstable question.
Synthesis quality depends on search, eligibility, coding, model choice, dependence among estimates and the quality of the underlying evidence.
Replication Changes Confidence
An intervention that works once under highly supported conditions is interesting. If independent teams reproduce benefits across multiple settings, confidence increases.
Failure to replicate does not automatically prove the original study was wrong. It may reveal implementation dependence, population differences, measurement change or ordinary sampling variation.
Evidence of No Effect Is Not the Same as No Evidence
A well-powered high-quality trial finding little effect can provide strong evidence that the intervention did not produce the expected outcome under those conditions. A small weak study with an uncertain estimate provides low evidence, not evidence of no effect.
Clearinghouse language should preserve this distinction.
Implementation Evidence Belongs Beside Effect Evidence
Suppose an intervention works when delivered by specialists with 40 hours of training and weekly coaching. A ministry considering national scale needs to know more than the effect size. It needs the delivery architecture.
- Who delivered it?
- How much training?
- How much time?
- What materials?
- What fidelity mattered?
- Which adaptations were allowed?
- What implementation failures occurred?
Cost Evidence Changes Feasibility
An intervention can have strong evidence of impact and still be unsuitable for scale because it requires scarce specialists or unsustainable recurrent spending.
Evidence infrastructure should link effect, evidence security, implementation and cost rather than force policy-makers to reconstruct the decision from separate documents.
External Validity Requires a Transfer Map
Before importing evidence, compare:
- learner age and prior attainment;
- subject and curriculum;
- teacher qualification and workload;
- class size and setting;
- language of instruction;
- delivery intensity;
- school resources;
- incentives and accountability;
- assessment outcome;
- time horizon.
The question is not “Was this evidence produced abroad?” The question is “Which causal conditions that mattered there are present here?”
Evidence Ratings Need Explanations
A star, tier or padlock can help a busy reader. It can also become a substitute for thinking. Every rating should be backed by inspectable reasons: number of studies, design strength, independence, consistency, precision, directness and other relevant limitations.
Do Not Collapse Different Outcome Domains
An intervention may improve attendance but not attainment, confidence but not completion, short-term test performance but not delayed retention. A clearinghouse should preserve outcome-specific evidence rather than issue one global label for the programme.
Practice Guides Add a Translation Layer
The What Works Clearinghouse publishes practice guides that combine research review with practitioner experience and expert judgement. This is valuable because evidence rarely arrives already formatted as a school routine.
The translation layer should state which parts are directly research-supported and which parts reflect expert synthesis or implementation judgement.
Policy Guidance Needs a Decision Context
“This intervention has strong evidence” is incomplete. Strong evidence for what outcome, in which populations, compared with what alternative, at what cost, over what duration?
Policy guidance becomes useful when it connects evidence to the actual decision.
Evidence Standards Need Version Control
Methods improve. Attrition rules, statistical practices, standards for clustered trials and synthesis techniques evolve. A study reviewed under one standard may receive a different judgement under a later version.
Store the standard version with every review rather than silently overwriting history.
Evidence Needs Freshness Rules
Some educational mechanisms are durable. Others depend on technology, curriculum or institutional context that changes quickly. A ten-year-old trial of one software environment may establish useful principles while saying little about a current product.
Freshness should be determined by what changed, not by age alone.
Commercial Evidence Needs Disclosure
Products can be effective. Commercial sponsorship does not invalidate evidence. It does create a conflict-of-interest field that belongs in the review.
Readers should know who funded the study, who analysed the data and whether independent replication exists.
Evidence Gaps Should Be Published
A clearinghouse should not only say what works. It should reveal where evidence is thin, where populations are missing and where important outcomes have not been measured.
Evidence gaps can guide research funding more rationally than another study in an already crowded area.
Negative Results Protect Public Money
Publishing only success creates a distorted market for education interventions. A credible evidence system makes null and negative findings discoverable so systems do not repeatedly purchase ideas that failed under similar conditions.
Uncertainty Should Survive Translation
Research reports often contain confidence intervals, caveats and subgroup limits. Policy summaries can accidentally remove them. The clearer the public guide becomes, the more deliberate the system must be about preserving what remains uncertain.
Case Study: The One Famous Study
Invented example: a ministry is attracted to a reading programme after one widely cited trial reports a large effect. Clearinghouse review finds high attrition, developer involvement and no independent replication.
The programme is not labelled ineffective. It is labelled promising but uncertain, and the ministry pilots it with independent evaluation before scale.
Case Study: The Small Effect With Strong Evidence
Invented example: an intervention produces a modest average improvement across twelve high-quality independent trials. The effect is smaller than a competing programme’s headline estimate but much more consistent.
The decision-maker can now compare magnitude, certainty, cost and implementation rather than being seduced by the largest number.
Case Study: The Evidence That Did Not Travel
Invented example: a mathematics intervention succeeds in small classes with specialist coaches. At national scale, ordinary teachers receive one day of training and no coaching. Effects disappear.
The clearinghouse had accurate impact evidence and insufficient implementation-transfer evidence. The next synthesis reports staffing and coaching as material conditions rather than side notes.
Case Study: The Meta-Analysis With Different Questions
Invented example: thirty studies labelled “digital tutoring” include adaptive practice, live human tutoring through video, homework platforms and AI feedback. The pooled average is precise and conceptually muddy.
The review is restructured around intervention mechanisms rather than product labels, and heterogeneity becomes interpretable.
Failure Modes and Repairs
- Effect size worship: repair by separating magnitude from evidence security.
- Design label worship: repair by reviewing execution, attrition, outcomes and analysis.
- Meta-analysis as automatic truth: repair by testing study comparability and quality.
- Developer evidence treated as universal: repair by recording independence and replication.
- Average effect hides heterogeneity: repair by analysing populations, settings and implementation conditions.
- Strong internal validity assumed to transfer: repair with a context transfer map.
- Evidence rating without reasons: repair with transparent review fields.
- Research effect without implementation: repair by publishing delivery requirements and fidelity evidence.
- Research effect without cost: repair by connecting evidence to resource implications.
- Policy summary deletes uncertainty: repair by preserving caveats in decision-facing products.
The Evidence Clearinghouse Operating Chain
- Define the decision question.
- Publish the review protocol.
- Define eligible populations, interventions, comparisons and outcomes.
- Define acceptable study designs by claim type.
- Search published and grey literature.
- Deduplicate records.
- Screen studies consistently.
- Apply evidence standards.
- Record conflicts of interest.
- Extract outcomes and implementation features.
- Record sample and context.
- Assess risk of bias.
- Assess independence and replication.
- Synthesise comparable evidence.
- Analyse heterogeneity.
- Estimate uncertainty.
- Rate evidence security.
- Separate outcome domains.
- Assess external validity.
- Assess implementation requirements.
- Connect cost evidence.
- Publish study-level decisions.
- Publish synthesis-level conclusions.
- Create practitioner or policy guidance.
- Label research-backed and judgement-based elements.
- Version standards and reviews.
- Monitor new evidence.
- Update conclusions when the evidence base materially changes.
- Publish unresolved evidence gaps.
An Evidence Infrastructure Dashboard
- review protocol version;
- standards version;
- studies screened;
- studies meeting standards;
- independent evaluations;
- replications;
- outcome domains;
- average effect estimate;
- uncertainty interval;
- heterogeneity;
- evidence-security rating;
- implementation requirements;
- cost evidence;
- populations represented;
- populations missing;
- contexts represented;
- conflicts of interest;
- last evidence search date;
- next review trigger.
Canonical Owner Boundaries
- Educational Research owns research design and methods broadly.
- Education Impact Evaluation & Causal Inference owns causal identification and impact evaluation.
- Education Research Governance, Ethics & Data Access owns ethics, participant protection and research-data access.
- Education Cost-Effectiveness & Benefit-Cost Analysis owns economic comparison of intervention costs and outcomes.
This node owns evidence qualification and translation: common research standards, clearinghouse review, synthesis, evidence-security ratings, replication, context transfer and practice-facing guidance.
The Return Path
Return to the policy meeting where everyone has a study.
The question is no longer who can cite research. The question is whether the system has a shared way to judge what each study can actually support.
Evidence standards make disagreement more productive because people can argue about explicit assumptions rather than trade impressive references.
The value of an evidence clearinghouse is not that it tells a system what to believe. It makes the reasons for belief inspectable.
Return to the How Education Works hub.