HSW-0248
You study two facts.
One is printed in large, clean type. The other is smaller and visually ordinary. The large item feels easier. It looks clearer. It may even feel more memorable.
Then the test tells you that font size did not meaningfully help your recall.
So surely your next prediction becomes more rational.
Not necessarily.
In a 2026 paper in Memory, Sofia Navarro-Báez, Arndt Bröder and Monika Undorf ran four experiments examining whether different kinds of feedback could repair metacognitive illusions in judgments of learning. Participants completed repeated study–test cycles, predicted future recall and received either no feedback or various forms of cognitive and metacognitive feedback. The researchers tested familiar illusions such as the tendency to judge large-font words as more memorable even when font size is not a useful memory cue, and the stability bias in which learners underappreciate how later study opportunities can change memory.
The striking result was not that people learned nothing from experience. Some judgments improved across cycles. The striking result was that feedback itself did not reliably repair the cue basis of the judgments. Cognitive feedback failed to eliminate several illusions. Additional metacognitive feedback partially reduced one stability bias in one experiment, but that result did not replicate in the next experiment.
Participants could sometimes state the correct general relationship after feedback—such as recognising that studying an item twice should improve recall—yet still fail to use that knowledge appropriately when predicting memory item by item.
HSW-0248 calls this problem metacognitive illusion repair: how a learner changes the cues used to judge “I know this,” “I will remember this,” or “this needs more study” after evidence shows that some cues are misleading.
The direct answer
Knowing that a cue is unreliable does not guarantee that the cue stops influencing your judgment.
Metacognitive monitoring is not a single verbal belief stored in one place. A learner can understand the rule “large print does not improve memory” and still experience large print as fluent, easy and familiar. That subjective experience can continue feeding item-level confidence.
This creates a practical distinction:
- Belief correction: “I now know this cue should not predict memory.”
- Judgment-policy correction: “When I estimate whether this specific item will be remembered, I actually stop relying on that cue and use better evidence.”
The first can happen without the second.
What is a metacognitive illusion?
A metacognitive illusion occurs when a judgment about one’s own cognition relies too strongly on information that is not valid for the outcome, or fails to rely enough on information that is valid.
In studying, common candidate cues include:
- how fluent an item feels;
- how familiar the page looks;
- whether the explanation was easy to read;
- how quickly an answer came to mind;
- how many times the item was studied;
- how visually distinctive the material is;
- whether a recent practice attempt succeeded.
Some of these cues can be genuinely informative in some contexts. The problem is not that subjective experience is always useless. The problem is assuming that because a cue feels diagnostic, it is diagnostic for the exact future performance that matters.
The font-size illusion is useful because the cue is so obviously superficial
Why study font size? Because it makes the logic visible. Large words often feel easier to process. That ease can raise judgments of learning even when actual recall gains are much smaller or absent.
In school, the misleading cue is rarely literally a font. It may be the neatness of notes, the recognisability of a textbook page, the smoothness of a teacher’s explanation, the fact that an answer “looks obvious” when the worked example is beside it, or the relief of recognising a multiple-choice option.
HSW-0248 does not claim those classroom examples are experimentally identical to the font-size illusion. They are analogies showing the same diagnostic question: Is the cue that makes learning feel good actually predicting independent later performance?
What the four experiments tested
Across the experiments, participants studied word lists over three study–test cycles. They made item-by-item judgments of learning, completed free-recall tests and then received different forms of feedback depending on the experimental group.
The feedback conditions included information about which items had been remembered, participants’ own earlier judgments, social-reference information in one experiment, and explicit metacognitive information designed to explain the relevant illusion in later experiments.
The researchers manipulated invalid or weak cues such as font size and font format alongside valid information such as whether an item would receive one or two study presentations. They then asked whether feedback would shift later judgments toward the valid cues and away from the invalid ones.
Cognitive feedback did not reliably repair the font-size illusion, the stability bias or the font-format illusion. Adding explicit metacognitive feedback partially remedied the stability bias in Experiment 3, but the improvement did not replicate in Experiment 4. Across several conditions, the influence of font size decreased with repeated task experience whether or not participants received feedback.
That final point matters. The result is not “people cannot improve their monitoring.” It is “the tested feedback formats did not reliably produce the intended correction beyond what repeated experience itself was already doing.”
Correct general knowledge can fail to control local judgments
In Experiments 3 and 4, many participants could answer follow-up questions correctly about which cues actually affected recall. They knew, at a general level, that twice-studied words were remembered better and increasingly recognised that large and small words could have similar recall.
Yet subgroup analyses did not show that this corrected knowledge reliably changed the cue basis of their item-by-item judgments.
This is a profound studying problem. A student may know:
- “Rereading can create familiarity.”
- “Recognition is easier than recall.”
- “A worked example beside me makes the problem easier.”
- “One good practice score can be noisy.”
And still behave as though familiarity, recognition, visible support or the last success were strong evidence of future independent performance.
Declarative advice is not automatically a new monitoring policy.
Illustrative case: the learner who knows the illusion and still trusts it
The following case is illustrative, not a participant from the experiments.
Arun has been taught about retrieval practice. He can explain why rereading can create an illusion of knowing. He tells his tutor, correctly, that being able to recognise a sentence is weaker evidence than being able to reconstruct it from memory.
That evening he rereads a biology chapter. Each paragraph feels familiar, so he marks almost every section “done.”
The next day, closed-book recall reveals large gaps.
Arun did not lack the metacognitive rule. He failed to activate and use it at the decision point. The fluent experience of reading still dominated the local judgment.
The repair is therefore not another lecture about the illusion. Arun needs a study process that forces cue validity into the judgment itself: predict first, retrieve, compare prediction with outcome, record which cue misled him, and repeat until the better cue becomes easier to use.
Feedback can tell you the result without changing what you attend to
Suppose a student predicts 90% recall and achieves 55%. Showing that mismatch is cognitive feedback. It tells the learner that the judgment was inaccurate.
But the learner may not know which cue caused the error. Was the material too familiar? Was the practice supported? Was the time delay underestimated? Was the successful example mistaken for a stable skill?
Even if feedback names the cue—“you overweighted ease of reading”—the cue can remain psychologically compelling on the next item. Repair requires changing cue use, not merely receiving a post-mortem explanation.
Experience can recalibrate even when formal feedback adds little
Across the 2026 experiments, some judgment accuracy improved over repeated cycles in control groups as well as feedback groups. Participants’ overall calibration shifted and the font-size effect on judgments sometimes weakened with task experience.
This suggests a useful distinction between being told and experiencing repeated cue–outcome mismatch. A learner who repeatedly predicts from fluency and then sees fluency fail may gradually reduce reliance on that cue, even if explicit feedback does not create an additional measurable benefit.
But experience alone is not guaranteed to educate. If the learner never records predictions, tests independently or notices which cue misled them, the mismatch can pass unnoticed. Repetition helps only when the relevant evidence is available to the learner’s control system.
A cue-validation protocol for studying
1. Name the cue behind the confidence
When you think “I know this,” ask what produced that feeling. Familiar page? Fast recognition? Recent repetition? Easy explanation? Successful prompted practice?
2. Make a prediction before the test
Prediction creates a record. Without one, post-test memory can rewrite how confident you remember being.
3. Test under the future performance condition
If you need free recall, test free recall. If you need to solve, solve. If you need to explain without notes, remove the notes. The outcome must correspond to the capability the judgment is supposed to predict.
4. Compare cue with outcome
Do not merely record “wrong.” Record the misleading cue: “felt easy because I had just read it,” “recognised the diagram but could not label it,” or “solved with formula sheet but failed without it.”
5. Install a replacement cue
Replace “looks familiar” with stronger evidence: successful delayed recall, correct explanation, independent solution, or repeated performance across varied examples.
6. Repeat across cycles
The goal is not one correct forecast. It is a monitoring policy that increasingly tracks the evidence that matters.
Why delayed judgments can help—but do not solve everything
One established route to better metacognitive monitoring is to delay the judgment so the learner must retrieve rather than simply inspect the just-studied item. eduKateSG’s How Studying Works | Delayed Judgments of Learning owns that timing mechanism.
HSW-0248 has a different job. It asks what happens when the learner already receives outcome information or an explicit warning about an illusion. The 2026 experiments show that corrective information can fail to alter the cue basis of later judgments.
The two ideas fit together: improve the evidence available at the moment of judgment, and train the learner to use that evidence repeatedly.
Monitoring can improve while memory itself stays unchanged
Metacognitive accuracy and memory performance are separate outcomes.
A learner can become better at predicting what will be remembered without learning more items. Conversely, a learner can improve recall through practice while remaining poor at identifying which items are fragile.
This matters for study control. Monitoring is useful because it helps allocate future effort. Better calibration should eventually improve decisions about what to restudy, retrieve, abandon or revisit—but that downstream benefit should be measured rather than assumed.
When measuring learning changes the learning
Asking learners to make judgments can itself alter behaviour. eduKateSG’s How Studying Works | Metacognitive Reactivity owns that question.
That means a cue-validation protocol should remain economical. The purpose is not to make students rate every fact forever. Early in training, explicit predictions can reveal bad cues. As monitoring improves, sampling can replace exhaustive rating.
The endpoint is not constant self-surveillance. It is faster, better-calibrated control.
Delayed and independent performance check
To test whether an illusion has actually been repaired, do not ask only whether the learner can recite the principle “fluency is misleading.”
Present a new set of items in which the misleading cue varies independently of real learnability. Ask for judgments before testing. Then examine whether confidence tracks the valid cue and outcome rather than the salient but invalid feature.
Next, change the surface cue. If the learner stopped trusting large font but immediately trusts colourful notes, the deeper monitoring policy may not have changed. The transfer test is whether the learner asks: “What evidence predicts future independent performance here?”
Finally, return after delay. A temporary correction during an explicit lesson is weaker evidence than spontaneous cue selection days later.
For parents: do not argue with confidence—ask what it is based on
When a child says, “I know it already,” asking “Are you sure?” often creates a confidence contest.
Ask instead: “What makes you think you know it?” If the answer is “because I just read it,” “because the notes look familiar,” or “because I got the same question right yesterday,” suggest one independent check.
The aim is not to make the learner distrust every feeling. It is to connect confidence to evidence that deserves confidence.
For tutors and teachers: feedback on accuracy is not the same as training cue use
If students repeatedly misjudge what they know, showing them the discrepancy is a start. But follow it with a cue question: “What information were you using when you made that prediction?”
Then arrange examples where the misleading cue and the valid cue come apart. A learner cannot learn cue validity if all easy-looking items are also easy to remember. Orthogonal variation makes the diagnostic relation visible.
Most importantly, require the learner to make the next judgment using the replacement cue. Correct knowledge that never controls the next decision remains inert.
Misconceptions
“Feedback cannot improve metacognition.” Too broad. The 2026 experiments tested particular feedback formats and particular JOL illusions. Some judgment measures improved with experience, and one feedback effect appeared in one experiment but did not replicate.
“People keep the wrong belief even after correction.” Not always. Participants could sometimes report the correct general cue relationship yet still fail to use it item by item. Belief and cue application can dissociate.
“Subjective fluency is useless.” No. Many subjective cues can be informative in some contexts. The problem is overweighting a cue when it does not predict the target performance.
“If calibration improves, learning improves.” Monitoring and memory are distinct. Better monitoring may support better study choices, but the downstream learning effect must still be demonstrated.
“Students just need to be told about illusions once.” The 2026 findings argue against assuming that explicit correction automatically changes item-level judgment policy.
Evidence and limits
The primary source is Navarro-Báez, Bröder and Undorf, “Mending metacognitive illusions in JOLs: when neither cognitive nor metacognitive feedback is effective”, published online 17 April 2026 in Memory and appearing in Volume 34, Issue 5.
Across four experiments, participants completed repeated study–test cycles with judgments of learning and different feedback conditions. Cognitive feedback did not repair the font-size illusion, stability bias or font-format illusion in the relevant experiments. Additional metacognitive feedback partially reduced the stability bias in Experiment 3, but the effect failed to replicate in Experiment 4. Some cue use and calibration changed across cycles independent of feedback condition.
The experiments used word-list learning and specific manipulated cues. They do not prove that teacher feedback, tutoring or every metacognitive intervention fails in classrooms. Nor do they establish that repeated experience is sufficient for all illusions. Their contribution is more precise: feedback about outcomes and even explicit information about a bias may fail to change the cues used in later item-level judgments.
The authors discuss a plausible mechanism: corrected beliefs may need to be activated at the moment of judgment to influence cue use. That explanation is theoretically grounded but should still be distinguished from the directly observed experimental outcomes.
The return
The hardest illusion to repair may be the one you can already explain.
You may know that familiarity is weak evidence. You may know that supported practice is easier than independent performance. You may know that one successful attempt does not guarantee retention.
Yet when the next item feels fluent, the old cue can still win.
So do not stop at the rule. Train the judgment.
Name the cue. Predict. Test. Compare. Replace weak evidence with stronger evidence. Repeat until the cue that deserves confidence increasingly becomes the cue that receives it.
That is the difference between knowing about metacognition and actually using it to steer study.
Continue through the How Studying Works Numbered Series Reading Index, or return to the How X Works Hub.