HSW-0130 · How Studying Works
A student studies twenty items.
After each one, we ask:
How likely are you to remember this later?
It sounds like measurement.
But the act of measuring can change the thing being measured.
Metacognitive reactivity occurs when making a judgment about learning changes attention, encoding, retrieval, strategy or later study decisions.
This matters because study systems often use confidence ratings, traffic-light colours, “know / unsure / don’t know” labels and predicted scores as if they were passive sensors. They are not always passive. Sometimes the rating itself becomes part of the learning event.
The 50-Second Read
- Monitoring can be reactive. A judgment of learning can alter memory or study behaviour.
- The effect is not uniformly positive. A 2025 meta-analysis found a small positive average effect, but outcomes varied with materials, relatedness and test conditions.
- Educational materials are not guaranteed to benefit. A 2026 study found no general improvement for several educationally relevant materials and negative reactivity in difficult conditions.
- Self-ratings can shift strategy. Recent research shows that repeated monitoring can push learners toward easier surface strategies in some rule-learning tasks.
- Ratings also change control. A 2026 study found that making judgments of learning increased later restudy choices.
- Measure deliberately. Ask a confidence question only when the answer will improve calibration or the next study decision.
- Keep performance receipts. Confidence remains a prediction; later retrieval and transfer are the evidence.
1. A Thermometer That Heats the Water
Good measurement is often imagined as invisible.
We observe the learner without changing the learner.
Metacognitive judgments complicate that picture. Asking someone to evaluate their own learning can change what they attend to and how they encode the material. It can trigger covert retrieval. It can alter motivation. It can change which items they choose to restudy.
So the study loop is not always:
learn → measure
It can become:
learn → make judgment → judgment changes learning → later performance
2. What Is a Judgment of Learning?
A judgment of learning, often abbreviated JOL in memory research, is a prediction about future remembering.
Examples:
- “80% chance I can recall this tomorrow.”
- “I know this well.”
- “Amber: I can recognise it but not explain it.”
- “I think I would score 7/10 on this topic.”
The lexical version of this problem already has a specialist owner in Vocabulary Metacognition and Judgments of Learning. This HSW article owns the broader study-system question: what changes when self-monitoring itself becomes part of study?
3. The Average Effect Is Small and Positive—But That Average Hides Important Variation
A 2025 meta-analytic review of immediate judgments of learning reported an overall improvement in memory performance of about g = 0.22. Related word pairs showed larger positive reactivity, while unrelated pairs showed a small negative effect. The authors also found signs of publication bias favouring positive results. See “Do immediate judgments of learning alter memory performance? A meta-analytical review”.
The educational lesson is not “add confidence ratings everywhere.”
It is:
making a judgment can change performance, and the direction depends on what the learner is studying and what the judgment causes them to do.
4. One Route: The Judgment Triggers Retrieval
To answer “Will I remember this?”, a learner may quietly try to retrieve the target.
That covert retrieval can itself strengthen memory.
A 2025 study reported evidence that judgments of learning can affect memory through covert retrieval. See Zhang and colleagues on JOL reactivity and covert retrieval.
This gives us a practical clue: a judgment may help when it forces the learner to inspect the cue-target relationship rather than simply state a feeling.
5. Another Route: The Judgment Changes Engagement
A 2024 EEG study found that making judgments of learning improved recognition performance in its task and was associated with greater learning engagement and elaborative processing. See Li and colleagues, 2024.
Again, the safe inference is conditional. The judgment may make the learner engage more deeply with some materials. That does not mean every rating scale improves every kind of learning.
6. The Measurement Can Also Make Learning Worse
Recent work has made the boundary clearer.
A 2025 study in Memory & Cognition found that judgments of learning could impair rule-based discovery when an easily memorised surface cue competed with a deeper relational rule. Repeated self-assessment appeared to encourage a conservative strategy shift toward the easier, immediately rewarded surface route. See “Judgments of learning impair rule-based discovery”.
This is a major warning for studying:
a monitoring prompt can make the learner optimise for what is easiest to judge rather than what is most important to learn.
7. Educational Material Does Not Automatically Show the Laboratory Benefit
A 2026 study tested immediate judgments of learning with key-term definitions, country outlines and animal species. The researchers found no general performance benefit for these educationally relevant materials and observed negative reactivity in some difficult study-test conditions. See “Immediate judgments of learning do not improve performance for educationally relevant materials”.
This is exactly why study advice should resist universal slogans.
A technique can have a real effect and still require careful scope.
8. The Rating Changes What the Learner Chooses to Study
Metacognitive reactivity is not limited to memory strength.
A 2026 study found that making judgments of learning increased the proportion of items people selected for restudy across several material types. See “Making Judgments of Learning Increases Restudy Choices”.
That means a confidence rating can alter the study portfolio itself.
If the judgment is badly calibrated, it can change not only what the learner believes but where the next thirty minutes go.
9. Mathematics: Confidence Can Shift Attention Toward the Answer Instead of the Structure
Imagine asking after every Mathematics example:
How likely are you to get the next one right?
That can be useful for calibration.
But if the learner starts optimising for immediate correctness, they may choose the most familiar procedure rather than inspect why the procedure applies.
A better rating can ask about the intended capability:
- Can I identify the method without a chapter label?
- Can I explain why it applies?
- Can I reject the nearest competing method?
- Can I do it after a delay?
The measurement should point toward the learning target, not merely the next answer.
10. English: “I Know This Word” Can Change How the Word Is Studied
Vocabulary judgments are especially revealing because a learner may use familiarity as the cue for “known.”
Once a word is marked green, it may disappear from review—even if productive use, collocation or register remains weak.
The specialist vocabulary owner should carry those lexical distinctions. The HSW principle is broader: classification changes allocation.
A self-rating system should therefore define what “green” means before the rating is collected.
11. Science: Confidence Can Attach to a Keyword Rather Than a Model
A student sees “diffusion” and feels highly confident.
The confidence may reflect recognition of the keyword rather than the ability to predict direction, explain particle behaviour, identify boundary conditions or apply the model in a new context.
Instead of asking “How well do you know diffusion?”, ask a more diagnostic prediction:
Could you explain what changes if the concentration difference becomes smaller, without looking?
The metacognitive question now points at the mechanism.
12. Calibration Still Matters
None of this means students should stop estimating what they know.
Study Calibration remains important because learners must compare prediction with later evidence.
The refinement is that the prediction itself is not inert.
A robust calibration loop is:
predict → perform → compare → update prediction rule → change study allocation → perform later under reduced support
13. Measure Less Often, but Better
If every card, paragraph and problem requires a confidence score, the monitoring burden itself can become study friction.
Use ratings when they answer a real decision:
- Which topic gets tomorrow’s review time?
- Which question should be attempted without notes?
- Where is confidence much higher than performance?
- Which concept appears weak but is actually stable?
A measurement that never changes a decision may not deserve to be collected.
14. The Delayed-Judgment Advantage
Immediate familiarity is a noisy cue.
One useful study design is to delay the judgment until after a brief retrieval attempt rather than asking for confidence while the answer remains visible.
The learner then bases the prediction on a more diagnostic event: what came back without the source.
This does not remove reactivity, but it aligns the judgment more closely with the capability being predicted.
15. The Financial Route: Measurement Changes Allocation
Study time is scarce capital.
A confidence label can increase or reduce the amount invested in a topic. That makes the label economically consequential even before any memory effect occurs.
Bad monitoring can misallocate study time. Reactive monitoring can additionally change the underlying asset being measured.
The study system therefore needs both:
- measurement quality;
- and awareness of measurement effects.
16. The School and System Route: Dashboards Are Not Neutral
Schools increasingly ask students to self-rate, reflect and track progress.
Those practices can be valuable. But repeated measurement can shape attention and behaviour, especially when the metric is tied to visible rewards, colour states or completion targets.
A student who is constantly asked “How confident are you?” may learn to optimise confidence management. A student asked “Which evidence changed your confidence?” is pushed toward a different kind of thinking.
Design the prompt around the educational behaviour you are willing to reinforce.
17. A Metacognitive Reactivity Audit
- Judgment: what exactly is the learner rating?
- Timing: before, during or after retrieval?
- Cue: what evidence will the learner use?
- Reactivity: how might the prompt alter attention or strategy?
- Control: what study decision changes because of the rating?
- Burden: how much time does monitoring consume?
- Calibration: is the prediction compared with later performance?
- Transfer: does confidence predict performance under changed conditions?
18. A 20-Minute Calibration Session
- 5 minutes: retrieve five items without notes.
- 2 minutes: rate confidence in delayed recall, not current familiarity.
- 5 minutes: study or repair weak items.
- 5 minutes: test the set again with changed order or cues.
- 3 minutes: compare predictions with outcomes and change tomorrow’s allocation.
19. What Not to Do
- Do not assume a self-rating is passive measurement.
- Do not collect confidence because a dashboard looks incomplete without it.
- Do not define “know” vaguely.
- Do not let confidence replace performance evidence.
- Do not ask so many monitoring questions that they fragment the learning task.
- Do not generalise word-pair findings mechanically to complex school learning.
- Do not interpret one study showing negative reactivity as evidence that metacognition is harmful.
20. Evidence Boundary
Judgment-of-learning reactivity is real, but heterogeneous. The 2025 meta-analysis finds a small positive average effect with substantial moderation; 2025 and 2026 work shows that prompts can produce null or negative effects in rule discovery and educationally relevant materials under some conditions. A judgment can alter memory, strategy and restudy control, but the direction is task-dependent.
The responsible educational use is therefore to treat metacognitive prompts as designed interventions whose benefits must be checked against their burden and behavioural side effects.
21. Return: When You Measure Learning, Watch What the Measurement Teaches
Metacognition is not a mirror floating outside the learning system.
The act of asking can change attention. The answer can change allocation. The rating can change strategy.
That does not make measurement useless.
It makes measurement part of the design.
Ask fewer, better questions—and compare every prediction with what the learner can actually do later.
Continue through Study Calibration, Learning Evidence Density, Study Metric Gaming, and the How Studying Works Numbered Series Reading Index.