A systematic review works by turning a broad question into an explicit evidence-finding and evidence-filtering process: define the question, pre-specify eligibility and methods, search widely enough to find the relevant evidence, distinguish studies from the papers that report them, screen records consistently, assess risk of bias, extract results carefully, synthesise what is sufficiently comparable, preserve what remains uncertain, and update the review when the evidence base changes.
People often imagine a systematic review as a very serious literature review.
That is too weak a description.
A traditional narrative review may begin with papers the author already knows, follow citations outward, emphasise studies that seem important and assemble an expert account.
A systematic review tries to make the route into the literature itself inspectable.
What counted as eligible?
Where did we search?
Which words did we use?
Which records were excluded?
Which apparently different papers actually came from the same study?
Which studies are biased enough that their results should be interpreted cautiously?
The governing question: if another careful team started with the same question and rules, could they reconstruct why this body of evidence—and not some convenient subset—became the basis of the conclusion?
Quick Read
QUESTION → PROTOCOL → ELIGIBILITY → SEARCH SOURCES → SEARCH STRATEGY → RECORDS → DEDUPLICATION → TITLE/ABSTRACT SCREENING → FULL-TEXT SCREENING → STUDY/REPORT LINKING → INCLUDED STUDIES → RISK OF BIAS → DATA EXTRACTION → EFFECTS / FINDINGS → SYNTHESIS → CERTAINTY → LIMITATIONS → CONCLUSION → UPDATE
Cochrane’s current Handbook describes systematic reviews as requiring a thorough, objective and reproducible search across relevant sources. Its 2025 Chapter 4 stresses that studies—not individual reports—are the units of interest, that multiple reports of the same study must be linked, that trial registries and other sources matter, and that search methods should be documented well enough to report and reproduce as far as possible.
1. The Review Begins With a Question, Not a Search Box
A weak review begins by typing a topic into Google Scholar.
A systematic review begins by deciding what evidence would answer the question.
For interventions, reviewers often use a framework such as PICO:
- Population or problem;
- Intervention;
- Comparator;
- Outcome.
Other review types may need exposure, context, diagnostic index test, reference standard, phenomenon of interest or qualitative setting.
The point is not the acronym.
The point is to define the evidence target before searching for evidence.
2. Scope Is a Design Decision
Too broad, and the review may combine questions that should remain separate.
Too narrow, and it may exclude relevant evidence or become useful only to a tiny niche.
Scope decides which populations, intervention variants, study designs, outcomes and time periods belong inside the review.
A good scope is broad enough to answer a meaningful question and narrow enough that included evidence remains interpretable together.
3. Eligibility Criteria Turn the Question Into Operational Rules
“Studies about vocabulary learning” is not an eligibility criterion.
Reviewers need rules that can be applied consistently:
- age or educational level;
- diagnosis or population characteristics;
- intervention definition;
- comparator definition;
- study design;
- minimum follow-up;
- outcome requirements;
- setting;
- publication type;
- language or date restrictions, if any.
Every restriction changes the evidence population.
4. Restrictions Need Justification Because They Can Create Selection Bias
English-only reviews are convenient.
They may miss eligible studies in other languages.
Published-journal-only reviews are convenient.
They may miss dissertations, regulatory data or trial-registry results.
A restriction is not automatically wrong.
It should be connected to a defensible reason and its likely consequence should be reported.
5. A Protocol Freezes the Main Rules Before the Literature Can Influence Them
Systematic reviews are vulnerable to the same retrospective flexibility as primary studies.
After reviewers see which studies exist, they may be tempted to adjust eligibility, outcomes or analysis rules.
A prospective protocol preserves the intended route.
Changes can still be made when justified, but the original and revised decisions remain distinguishable.
This is the review-level expression of How Preregistration Works.
6. Search Strategy Is an Information-Retrieval Problem
The search has two competing goals.
Sensitivity: find as many relevant reports as possible.
Precision: avoid retrieving overwhelming numbers of irrelevant reports.
Cochrane recommends maximising sensitivity while striving for reasonable precision.
Missing a relevant study can bias a review.
Retrieving extra irrelevant records mainly costs screening time.
That asymmetry often pushes systematic searches toward recall rather than convenience.
7. Search Concepts Are Not the Same as the Full Research Question
Trying to search every PICO element can reduce sensitivity.
Outcomes and comparators may be poorly represented in titles, abstracts or indexing.
Cochrane notes that intervention searches often focus on a smaller set of central concepts such as condition, intervention and eligible study design.
The search query is a retrieval instrument, not a prose restatement of the question.
8. Controlled Vocabulary and Free Text Catch Different Language
Databases such as MEDLINE assign controlled subject headings.
Authors also use ordinary words, abbreviations, older terminology, brand names and spelling variants.
Strong searches typically combine both controlled vocabulary and free-text terms.
One catches the database’s classification.
The other catches the authors’ language.
9. OR Broadens Within a Concept
Suppose the review concerns adolescents.
Relevant papers might use adolescent, teenager, youth, secondary school student or other terms.
These synonyms and related representations are usually linked with OR.
OR tells the retrieval system that any of these may represent the concept.
10. AND Connects Different Concepts
Once the population concept is broad, it can be connected with an intervention concept using AND.
AND narrows the result to records containing evidence of both conceptual regions.
Boolean logic sounds mechanical, but the choice of concepts determines which literature becomes visible.
11. One Database Is Rarely the Whole Literature
Cochrane’s 2025 guidance explicitly states that MEDLINE alone is inadequate for intervention reviews.
Different databases cover different journals, disciplines, countries and publication types.
An education review may need ERIC, PsycINFO, Scopus, Web of Science or discipline-specific databases.
A biomedical review may use CENTRAL, MEDLINE and Embase plus specialist sources.
Database selection should follow the evidence ecology of the question.
12. Trial Registries Reveal Studies That Journals May Not
A trial can exist without a journal article.
Its results may be delayed, unpublished or available only in a registry.
Searching registries helps reviewers find studies independent of publication success.
This is one defence against publication bias and missing evidence.
13. Regulatory Sources Can Contain Evidence Missing From Journals
Clinical study reports and regulatory documents can contain detailed outcomes, harms and analyses not fully represented in journal publications.
Cochrane increasingly emphasises regulatory sources in comprehensive searching.
The visible scholarly literature is not always the complete evidence record.
14. Citation Searching Moves Through the Literature Graph
Database searches find records through indexed fields.
Citation searching uses relationships among papers.
Backward citation searching asks what an included paper cited.
Forward citation searching asks which later papers cited it.
These routes can recover studies missed by terminology or indexing.
15. Search Strategies Benefit From Independent Peer Review
A missing synonym, incorrect Boolean bracket or inappropriate filter can silently remove relevant evidence.
Cochrane strongly recommends peer review of search strategies.
Information specialists and librarians are therefore methodological collaborators, not merely database operators.
16. Search Date Is Part of the Evidence Boundary
A systematic review published in August may have a search ending in January.
The conclusion therefore represents evidence available up to the search date, not automatically the publication date.
Readers should look for the last search date because new trials can make an apparently recent review scientifically stale.
17. Search Alerts and Updating Turn Reviews Into Maintained Evidence Products
Some questions evolve quickly.
A static review can become obsolete within months.
Living systematic reviews use repeated searches and planned update mechanisms to incorporate important new evidence.
The review becomes a maintained knowledge system rather than a one-time publication.
18. Search Results Are Records, Not Studies
This distinction is foundational.
A database record usually points to a report.
A report describes all or part of a study.
One study may generate:
- a protocol;
- a conference abstract;
- a main results paper;
- a harms paper;
- a subgroup paper;
- a long-term follow-up paper;
- a registry record.
Counting these as seven independent studies would be a major error.
19. Deduplication Removes Duplicate Records of the Same Report
The same article may appear in MEDLINE, Embase, Scopus and Web of Science.
Reference managers and review software identify duplicate records.
Deduplication is not the same as study linking.
Deduplication says two database records point to the same report.
Study linking says different reports came from the same underlying experiment or cohort.
20. Title and Abstract Screening Is Designed to Be Over-Inclusive
At the first screening stage, reviewers usually remove records that are clearly irrelevant.
Ambiguous records are retained for full-text assessment.
Cochrane recommends being generally over-inclusive at this stage.
A false inclusion costs time later.
A false exclusion can permanently remove eligible evidence.
21. Full-Text Screening Applies the Actual Eligibility Rules
A title can look eligible while the full study is not.
The wrong age group.
The wrong comparator.
No required outcome.
A non-randomised design where only randomised trials were eligible.
Full-text screening converts the abstract-level candidate pool into the actual study set.
22. Exclusion Reasons Preserve the Boundary
Readers often know of a famous study and wonder why it is absent.
Systematic reviews can list plausibly eligible excluded studies and state the primary reason for exclusion.
This turns absence into an auditable decision rather than a silent disappearance.
23. Independent Screening Reduces Idiosyncratic Decisions
Eligibility criteria still require judgement.
Two reviewers can screen independently, compare decisions and resolve disagreements through discussion or a third reviewer.
The value is not that two people are automatically right.
It is that disagreement becomes visible and arbitrary interpretation becomes harder to hide.
24. Automation Can Assist Screening Without Owning the Inclusion Decision Blindly
Machine-learning tools can rank records, identify likely duplicates and reduce repetitive work.
But automation creates its own recall risk.
If a classifier systematically misses unusual terminology, the review can inherit that blind spot.
Automation should therefore be validated against the review’s required sensitivity and documented clearly.
25. PRISMA Makes the Selection Route Visible
PRISMA reporting typically shows how many records were identified, deduplicated, screened, excluded and included.
The flow diagram is a receipt for the funnel from a large retrieval set to the final evidence set.
PRISMA does not guarantee that the search was comprehensive or the judgements were good.
It makes the process easier to inspect.
26. Risk of Bias Is Different From Study Quality as a Single Score
Older reviews sometimes gave each study a numerical quality score.
Modern risk-of-bias approaches prefer domain-specific judgement.
Was allocation random?
Was allocation concealed?
Could missing outcomes distort the estimate?
Was outcome measurement influenced by knowledge of intervention?
Were results selectively reported?
Different biases have different causal routes into the estimate.
27. Risk of Bias Is About the Result, Not Moral Character
A study can be conducted by careful, honest researchers and still have high risk of bias.
A study can also have low risk of bias in one outcome and higher risk in another.
The question is not whether the researchers are good people.
The question is whether the design and conduct create a plausible route by which the estimated result could be systematically distorted.
28. Data Extraction Is Another Measurement Stage
Reviewers extract sample sizes, participant characteristics, intervention details, outcome definitions, follow-up times, effect estimates and uncertainty.
Extraction errors can change the final synthesis.
Independent extraction or verification reduces transcription errors and inconsistent interpretation.
A review is only as accurate as the evidence representation it builds from the source reports.
29. Study Characteristics Are Needed to Judge Comparability
Two trials can carry the same intervention label while differing in dose, duration, teacher training, participant age, baseline risk or implementation quality.
Extraction tables make these differences visible.
Without them, the review may collapse distinct interventions into one name.
30. Outcome Harmonisation Is Often Harder Than It Looks
One study reports reading comprehension at six weeks.
Another reports vocabulary at one day.
A third reports transfer at six months.
All are educational outcomes.
They are not necessarily the same construct or time horizon.
Systematic synthesis requires enough conceptual separation to avoid averaging unlike outcomes merely because they are numerical.
31. Narrative Synthesis Is Not a Failed Meta-Analysis
Sometimes studies are too heterogeneous for meaningful quantitative pooling.
A structured narrative synthesis can compare study direction, magnitude, context, bias and mechanisms without forcing a single number.
The absence of a pooled estimate can be a sign of methodological restraint.
Systematic review owns the evidence synthesis whether or not meta-analysis is appropriate.
32. Meta-Analysis Is Optional, Not Mandatory
When effect estimates are sufficiently compatible, meta-analysis can increase precision and explore heterogeneity.
When they are not, pooling can manufacture an answer that corresponds to no coherent scientific question.
The distinction is simple:
Systematic review decides what the evidence base is and how it should be appraised. Meta-analysis is one possible statistical operation on a subset of that evidence.
33. Certainty of Evidence Is Not the Same as Number of Studies
Twenty small biased studies do not automatically provide high-certainty evidence.
Certainty frameworks such as GRADE consider dimensions including risk of bias, inconsistency, indirectness, imprecision and publication bias.
The purpose is to distinguish a precise estimate from a trustworthy estimate.
34. Indirectness Asks Whether the Evidence Matches the Receiver’s Question
A trial in adults may not answer a question about children.
A six-week surrogate outcome may not answer a question about long-term function.
An intervention delivered by specialists may not represent ordinary schools.
Indirect evidence can still be informative.
Its distance from the actual decision should remain visible.
35. Imprecision Asks Whether Important Alternatives Remain Compatible With the Evidence
A review may estimate benefit but have a confidence interval that includes substantial benefit, negligible effect and modest harm.
The average direction is not enough.
Decision-making depends on which meaningful possibilities remain plausible.
This connects directly to How Statistical Power Works and the confidence-interval layer that follows it.
36. Inconsistency Can Be a Discovery, Not Just a Downgrade
If studies disagree, reviewers should ask why.
Different intervention intensity?
Different populations?
Different follow-up?
Different bias?
A heterogeneous literature can reveal where an effect works and where it fails.
37. Missing Evidence Can Enter Before and After Publication
An entire study may never be published.
A published study may omit one unfavourable outcome.
A conference result may never become a full paper.
A registry may contain results absent from the article.
Systematic reviews try to reconstruct these missing regions rather than equating “published papers” with “all evidence”.
38. Systematic Review Does Not Mean Bias-Free Review
A protocol can be poorly designed.
A search can miss databases.
Reviewers can apply eligibility inconsistently.
Risk-of-bias assessment can be superficial.
Meta-analysis can pool incompatible studies.
Systematic means the process is explicit and structured.
Quality still depends on the decisions inside that structure.
39. A Review of Reviews Adds Another Layer of Compression
Umbrella reviews or overviews synthesise systematic reviews rather than primary studies.
This can cover broad evidence landscapes efficiently.
It introduces new dependence problems because multiple reviews may include many of the same primary studies.
Compression layers inherit overlap and bias from layers below.
40. Scoping Reviews Ask a Different Question
A scoping review often maps what evidence exists, how a field is defined, which methods are used and where gaps remain.
It may not be designed to estimate one intervention effect.
Calling every structured literature search a systematic review erases useful distinctions among review designs.
41. Rapid Reviews Trade Comprehensiveness for Timeliness
Decision-makers sometimes need an answer quickly.
Rapid reviews streamline parts of the systematic-review process: fewer databases, narrower searching, single-reviewer steps with verification or other shortcuts.
The trade-off should be explicit.
Speed is a design choice with possible coverage costs.
42. Living Reviews Treat Staleness as a Failure Mode
In fast-moving domains, the important review question is not only “Was this correct when published?”
It is also “Is it still current enough to guide a decision?”
Living reviews repeatedly search, screen and integrate new evidence according to predefined update rules.
Maintenance becomes part of scientific validity.
43. Reviews Can Be Wrong Because the Underlying Literature Is Wrong
A systematic review cannot repair fabricated data it fails to detect.
It cannot recover outcomes never measured.
It cannot transform severe confounding into randomisation.
It can identify, classify and down-weight confidence in flawed evidence.
It cannot manufacture evidence that was never created.
44. Retractions and Corrections Matter to Review Maintenance
A study included today can be corrected or retracted tomorrow.
Search and update processes should notice major integrity changes.
Cochrane’s current search guidance explicitly discusses identifying fraudulent studies, retractions, errata and comments.
The evidence estate has a lifecycle.
45. Education Systematic Reviews Need Strong Intervention Identity
“Small-group tuition” can mean two students or twelve.
It can mean diagnostic teaching or worksheet supervision.
It can mean thirty minutes weekly or six hours weekly.
If reviewers pool only labels, educational mechanisms disappear.
Intervention coding should preserve the active ingredients relevant to learning.
46. Technology Reviews Need Version Awareness
A study of an AI system from 2022 may not represent the system available in 2026.
Software, models, hardware and interfaces change.
Systematic reviews of rapidly evolving technology should preserve product version, date and deployment context rather than treating a brand name as a timeless intervention.
47. The Hostile Test: The Review That Searches Only One Database
A reviewer searches one convenient database and finds twelve positive trials.
Unpublished registries contain four negative trials.
Another database contains three small null trials from a different region.
The review may be internally tidy and externally incomplete.
Systematic searching exists because convenience can create a biased evidence sample.
48. The Second Hostile Test: Counting Papers Instead of Studies
One major trial produces a protocol, main paper, subgroup paper and long-term follow-up.
The review counts all four as independent studies.
That trial now receives artificial conceptual weight even before any meta-analysis occurs.
The unit of evidence is the underlying study, not the publication count.
49. The Third Hostile Test: Eligibility Changes After Seeing Results
The protocol includes follow-up of at least one month.
Several favourable two-week studies appear.
The reviewers quietly remove the minimum follow-up rule.
The change may or may not be scientifically defensible.
The problem is invisible revision.
Protocol deviation should be documented and justified.
50. The Fourth Hostile Test: Meta-Analysis Because Software Offers a Button
The included studies measure different constructs, time points and intervention versions.
The software converts everything to standardised mean differences.
A pooled number appears.
Mathematical compatibility is not conceptual compatibility.
The correct systematic review may be a structured synthesis without a pooled estimate.
51. The Fifth Hostile Test: A Beautiful PRISMA Diagram Hiding a Weak Search
The review reports every screening number perfectly.
The search strategy omitted the main synonym used by half the field.
Transparent reporting reveals the process.
It does not guarantee that the process was adequate.
52. Primary School: Systematic Review Begins as “Look for All the Fair Evidence”
A child asks whether plants grow better in sunlight.
They find three classroom experiments.
The first supports sunlight strongly.
The second is unclear.
The third disagrees.
The child should not hide the inconvenient experiment.
The early habit is simple:
Find the relevant evidence first. Decide what it means second.
53. Secondary School: Separate Search, Selection and Conclusion
Students can learn a three-stage discipline:
- Define what counts as relevant.
- Search and record all matching evidence.
- Only then compare quality and conclusion.
This prevents conclusion-first research where only supporting sources are collected.
54. JC and University: Systematic Review Becomes Evidence Provenance
At higher levels, a learner should be able to reconstruct:
- question;
- protocol;
- eligibility;
- databases and registries;
- full search strings;
- search dates;
- deduplication;
- screening process;
- study/report linking;
- risk of bias;
- extraction;
- synthesis choice;
- certainty judgement;
- deviations;
- update status.
The review becomes a documented transformation from the world’s available evidence to one bounded conclusion.
55. Where Systematic Reviews Fit in the eduKateSG “How Works” Landscape
- How Scientific Research Works — primary research feeding cumulative knowledge.
- How Preregistration Works — protocol and decision timing.
- How Sampling Works — study selection as a sample of the evidence universe.
- How Research Bias Works — bias in primary studies and review selection.
- How Meta-Analysis Works — optional quantitative synthesis.
- How Peer Review Works — scrutiny of the review and its protocol.
- How Search Works — the general retrieval mechanism beneath database searching.
- How Citation Works — the literature graph used for forward and backward discovery.
Systematic review owns a distinct canonical job: construct the most defensible available evidence set for a defined question and make the route from search to conclusion inspectable.
56. What This Article Does Not Claim
- A systematic review is not automatically high quality merely because it follows a reporting checklist.
- A systematic review does not require meta-analysis.
- A meta-analysis is not a substitute for a systematic search and explicit eligibility process.
- One database is rarely sufficient for comprehensive evidence retrieval.
- Reports are not the same units as studies.
- More included studies do not automatically mean higher-certainty evidence.
- Risk of bias is not a moral judgement about researchers.
- A protocol can be changed when justified, but major deviations should remain visible.
- A review cannot repair evidence that was never measured or data that are fundamentally invalid.
- A recent publication date does not guarantee a recent search date.
57. A Compact Systematic Review Audit
- What exact question does the review answer?
- Was a protocol written prospectively?
- What eligibility criteria were defined?
- Are language, date or publication restrictions justified?
- Which databases were searched?
- Were trial registries searched?
- Were regulatory or grey-literature sources relevant?
- Were citation searches used?
- Was an information specialist involved?
- Was the search strategy peer reviewed?
- When was the last search run?
- Are full search strings available?
- Were records deduplicated?
- Were studies distinguished from reports?
- How was title/abstract screening performed?
- How was full-text screening performed?
- Were disagreements resolved independently?
- Are exclusion reasons available?
- Was risk of bias assessed by domain?
- Was data extraction verified?
- Are intervention and outcome definitions comparable?
- Was meta-analysis scientifically justified?
- If no meta-analysis was done, was narrative synthesis structured?
- Was certainty of evidence assessed?
- Could missing evidence affect the conclusion?
- Were protocol deviations reported?
- Does the search date make the review current enough for the decision?
- Is an update mechanism needed?
58. Frequently Asked Questions
What is a systematic review?
A systematic review is a research process that uses explicit, pre-defined methods to identify, select, appraise and synthesise all eligible evidence relevant to a clearly defined question as comprehensively and transparently as feasible.
What is the difference between a systematic review and a meta-analysis?
A systematic review constructs and appraises the evidence base. Meta-analysis is an optional statistical method for combining compatible effect estimates from some or all included studies.
Why search more than one database?
Different databases index different journals and disciplines. Searching more sources reduces the risk that the evidence set is biased by the coverage of one database.
Why are studies different from reports?
One underlying study can generate several publications, abstracts, registry entries and follow-up reports. Systematic reviews must link them so the study is not accidentally counted multiple times.
Can systematic reviews become outdated?
Yes. A review represents evidence available up to its last search. Rapidly evolving topics may require frequent updates or a living-review model.
59. Authoritative Research Corridor
- Cochrane Handbook Chapter 4 — Searching for and Selecting Studies, updated March 2025
- Cochrane Handbook Chapter 3 — Defining Criteria for Including Studies and Grouping for Synthesis
- Cochrane Handbook Chapter 8 — Assessing Risk of Bias in a Randomized Trial
- PRISMA 2020 Statement — Reporting Systematic Reviews
- Cochrane Handbook Chapter 10 — Meta-Analysis
Final Thought: A Systematic Review Is a Controlled Journey Through What We Know
Every literature contains famous papers.
Every researcher has favourite papers.
Every search engine has ranking biases.
Every journal system makes some results easier to find than others.
A systematic review does not escape those forces by declaring itself systematic.
It responds by exposing the route.
This is the question.
These were the rules.
These were the databases.
These were the search terms.
These records were excluded for these reasons.
These reports belonged to these studies.
These studies carried these risks of bias.
This evidence could be combined.
This evidence could not.
And this is how much confidence the remaining evidence deserves.
The power of a systematic review is not that it knows every paper. It is that it tries to make the path from an enormous, uneven literature to one defensible evidence set visible enough to challenge, reproduce and improve.