HSW-0261 · How Studying Works
Retrieval practice has one of the strongest reputations in learning science.
That reputation is deserved.
Across many experiments and classrooms, trying to bring information back from memory often produces better later retention than simply seeing the information again.
But a strong average result can quietly turn into a stronger claim than the evidence supports:
Retrieval practice works the same way for everyone.
That conclusion does not follow automatically.
Generalisability is a separate scientific question from average efficacy. A method can show a reliable mean advantage across many studies while important learner groups, task conditions or accessibility needs remain underrepresented, insufficiently compared or dependent on adaptations that the headline effect does not specify.
This article owns that boundary. The broad mechanism remains with How Retrieval Practice Works. Here we ask what learners, parents and tutors should conclude—and what they should not conclude—when a robust learning principle meets a heterogeneous population.
Quick Answer
Retrieval practice has substantial evidence behind it. The correct response to uncertainty about learner differences is not to abandon retrieval. It is to distinguish three claims:
- Average efficacy: does retrieval practice outperform a comparison condition on average?
- Heterogeneity: does the size or form of the benefit differ across learners or contexts?
- Accessibility: what supports are required so the target learner can make a meaningful retrieval attempt?
A 2026 npj Science of Learning perspective examined the testing-effect literature and argued that the field has increasingly moved toward classroom application while still giving limited attention to some learner subpopulations, using dyslexia as a case study. Importantly, the paper does not demonstrate that retrieval practice fails for dyslexic learners. It identifies a coverage and generalisation problem that needs direct empirical testing. See Wilschut, Sense and van Rijn, 2026.
A research gap is not evidence of failure. It is evidence that confidence should stop where direct evidence stops.
1. Why Average Effects Are So Useful—and So Easy to Overread
Suppose a large collection of studies finds that retrieval practice improves later memory compared with restudy.
That result is valuable. It tells us something important about learning under the conditions represented in those studies.
But an average combines many things:
- learners with different prior knowledge;
- different ages;
- different subjects;
- different retrieval formats;
- different feedback schedules;
- different retention intervals;
- different levels of retrieval success;
- different language, reading and response demands.
The average effect is not a description of one universal learner. It is a summary over variation.
2. The Testing Effect Is Robust, but Robust Does Not Mean Context-Free
Meta-analytic evidence has repeatedly found a meaningful testing effect across many conditions. For example, Rowland’s 2014 meta-analysis reported a substantial overall advantage for practice testing over restudy while also showing that effect size varied with test format, feedback, retention interval and other design features.
That is exactly what a mature learning principle should look like: reliable enough to use, conditional enough to design carefully.
“Robust” should mean survives many reasonable changes. It should not mean immune to every change.
3. The 2026 Coverage Audit
Wilschut and colleagues analysed publication trends in testing-effect research and found a clear shift from basic laboratory work toward applied educational questions.
That is encouraging. It means the field has increasingly tested retrieval in the places where students actually learn.
At the same time, the authors argued that research attention has not been evenly distributed across learner populations. Their perspective uses specific learning disabilities—and dyslexia in particular—as a case where theoretical reasons exist to ask whether standard retrieval tasks, especially those with substantial reading or verbal-output demands, should simply be assumed to operate identically.
The scientific job is therefore not to announce a hidden exception. It is to design studies capable of finding one if it exists.
4. Separate the Memory Operation From the Response Barrier
A retrieval task may ask the learner to remember a concept, but the visible response can also require:
- rapid reading;
- spelling;
- handwriting;
- typing;
- oral production;
- visual discrimination;
- sustained attention to a long prompt;
- holding several instructions in mind.
If a learner performs poorly, the failure may occur at memory retrieval, response production, access to the cue, or some combination.
This is why retrieval practice should preserve the target cognitive operation while allowing irrelevant barriers to be reduced when appropriate.
5. Accessibility Is Not the Opposite of Retrieval
There is a mistaken idea that retrieval practice must always mean “no help at all”.
A learner can retrieve from:
- a spoken cue instead of a dense written cue;
- a diagram label instead of a paragraph;
- a first-letter prompt;
- a partially completed representation;
- a concept question with simplified wording;
- an oral response instead of handwriting, when handwriting is not the target.
The support changes the retrieval demand. That must be acknowledged. But a supported retrieval attempt can still require the learner to bring target knowledge to mind.
The correct design question is: which support removes an irrelevant access cost without supplying the answer the learner is meant to retrieve?
6. Retrieval Success Is Not Just an Outcome; It Changes the Treatment
Many accounts of retrieval practice assume that successful retrieval contributes to later learning.
If one group succeeds on 80% of practice trials and another succeeds on 25%, the two groups have not actually received the same cognitive experience even if the worksheet was identical.
This is a general reason to monitor:
- practice-test success;
- error type;
- feedback uptake;
- time to response;
- support required;
- delayed retention.
A method can be nominally identical and functionally different.
7. Mathematics Example: Retrieval Without Reading Noise
Imagine the target is whether a student can retrieve the relationship between gradient, vertical change and horizontal change.
A dense paragraph may make reading part of the task. If reading comprehension is not the target, a tutor can ask orally:
“What does gradient compare?”
The student still has to retrieve the mathematical relation.
Later, if the examination uses written questions, reading demands must also be practised. But they do not need to contaminate every early retrieval opportunity.
8. English Example: Retrieval May Be the Skill
In vocabulary learning, spelling or exact wording may sometimes be part of the target capability.
If the goal is productive written vocabulary, accepting only an oral recognition response would undertrain the final performance.
This is why accessibility cannot be defined as making every task easier. It means aligning support with the construct being learned.
9. Science Example: Let the Learner Retrieve the Mechanism in More Than One Form
A learner can retrieve a mechanism through a diagram, oral explanation, written paragraph or prediction.
If all four formats point to the same underlying model, varying the response format can help distinguish weak scientific knowledge from one narrow performance bottleneck.
The final test should still include the response form the real task requires.
10. A Generalisability Audit for Retrieval Practice
Before declaring retrieval “not working” for a learner, inspect the system.
- Target: What exact knowledge or capability should be retrieved?
- Cue: Can the learner access and understand the prompt?
- Response: Does the response format add unrelated difficulty?
- Success: Is retrieval occurring often enough to provide a learning event?
- Feedback: Is correction immediate enough and clear enough to repair errors?
- Delay: Does the knowledge survive after freshness fades?
- Transfer: Can it be retrieved with a changed cue or in a fresh problem?
This audit protects against two opposite errors: abandoning a strong method too quickly, and forcing one implementation regardless of evidence.
11. What Heterogeneity Can Mean
When researchers say an effect may be heterogeneous, several different realities are possible:
- the underlying memory benefit differs;
- baseline performance differs but proportional benefit is similar;
- one group needs more feedback;
- one group receives fewer successful retrieval events;
- the standard task includes an access barrier;
- measurement is noisier in one group;
- the effect is actually similar and the apparent difference is sampling variation.
These possibilities require different responses. “Different learner” is not a mechanism.
12. The Danger of Personalisation by Stereotype
Once learner differences are discussed, a new error becomes possible: assuming that membership in a category predicts the best study method for an individual.
Do not replace one universal rule with another.
Instead, use individual performance evidence:
- Does the learner understand the cue?
- Can they retrieve with support?
- Does reducing support improve or collapse performance?
- Which feedback repairs the error?
- Does improvement survive a delay?
Adaptation should be evidence-responsive, not stereotype-responsive.
13. The Retrieval Ladder
When unsupported recall is failing, a learner can move through a graded sequence:
- free recall;
- topic cue;
- semantic hint;
- partial answer;
- recognition among alternatives;
- corrective restudy;
- return immediately to recall.
The point is not to reach the bottom of the ladder. The point is to use the smallest support that restores a productive retrieval event, then move upward again.
14. Retrieval Practice vs Testing Pressure
Low-stakes retrieval practice is not the same as high-stakes assessment.
If every retrieval attempt feels like a public judgement, the task changes motivationally and socially.
For learning, retrieval often works best when errors are usable information rather than identity evidence.
Students can still be held to rigorous standards while practice remains safe enough to expose incomplete knowledge honestly.
15. Retrieval Practice vs the Retrieval-Format Ladder
Retrieval Format owns the distinction among recognition, cued recall and free recall.
Generalisability asks another question: when a format is chosen, how confident are we that the observed benefit and burden apply across learners?
16. Retrieval Practice vs Challenge Point
The Challenge Point owns the broad principle that useful difficulty depends on learner capability.
Generalisability adds an evidence boundary: researchers should not infer that one implementation produces identical benefit simply because the average effect is strong.
17. For Researchers: Report the Distribution, Not Only the Mean
Average treatment effects remain important.
But educational interpretation improves when studies also report:
- baseline performance;
- retrieval success during practice;
- variability;
- dropout and missingness;
- accessibility conditions;
- subgroup sample sizes;
- task and response demands.
That makes it easier to tell whether a method failed, the implementation failed, or the study was never designed to answer the subgroup question.
18. For Schools: Standardise the Principle, Adapt the Access Route
A school can standardise a principle such as “students regularly retrieve important knowledge” without forcing every student through one identical prompt and response format.
That distinction allows consistency without pretending uniformity.
The standard can be:
- knowledge must be brought back without simply rereading;
- feedback must repair errors;
- support should be reduced toward the target performance;
- retention must be checked after delay.
The exact route can vary.
19. Parent and Tutor Guide
If retrieval practice appears unusually difficult, do not immediately conclude that the learner “cannot learn by testing”.
Run four comparisons:
- same knowledge, simpler cue;
- same knowledge, different response mode;
- same knowledge, immediate corrective feedback;
- same knowledge, delayed retest.
If performance improves, the original task may have bundled retrieval with an avoidable barrier.
If performance remains weak across formats, the learner may need stronger initial learning, prerequisite repair or more supported retrieval.
20. Evidence Boundary
The 2026 Wilschut, Sense and van Rijn paper is a perspective and bibliometric coverage analysis. It does not experimentally compare retrieval-practice effects across all neurodiverse populations, and it does not establish that dyslexia or another learning difference removes the testing effect.
Its contribution is a warning about external validity: widespread educational adoption should be accompanied by direct evidence in the learners who will actually receive the intervention.
Educationally, that means preserving a strong evidence-based method while measuring whether the implementation produces successful, accessible retrieval and durable learning for the learner in front of us.
21. Return: Strong Evidence Deserves Strong Boundaries
Retrieval practice does not become weaker science when we ask where its evidence is strongest.
It becomes better science.
Use the robust average effect as a starting prior. Then inspect retrieval success, access barriers, feedback, delay and transfer. Adapt the implementation when evidence says the learner needs it. And never turn “not yet studied enough” into “does not work”.
Continue through How Retrieval Practice Works, Retrieval Format, The Challenge Point, the How Studying Works Numbered Series Reading Index and the How X Works Hub.