HSW-0213 · How Studying Works
A student can become more reflective about learning and, at the same time, become worse at discovering what the problem is really about.
That sounds contradictory because metacognition is usually treated as an unquestioned good. Students should monitor themselves. They should ask whether they understand. They should rate confidence, inspect progress and decide what to do next. Those habits can be valuable. But the act of monitoring is not always neutral. Sometimes the question you ask about learning changes the way the learner learns.
Quick answer
A 2025 Memory & Cognition study found that asking participants for judgments of learning could alter category-learning strategy. When both a deeper relational rule and more obvious visual features could be used to classify examples, participants asked to judge their learning were worse at discovering the relational rule, while feature memorisation was not harmed. When the task was redesigned so that only the relational rule was useful, the researchers did not find the same negative reactivity.
The important lesson is not “stop metacognition.” It is narrower: measurement can redirect attention and strategy. If a learner is repeatedly asked, “How well do you know this?” the easiest evidence available may be surface familiarity, visible features or recent success. Those cues can become the route the learner optimises, even when the real educational goal is to discover a relation that survives new examples.
This article owns one job: how to monitor learning without allowing the monitoring prompt to quietly replace the deeper learning objective.
A judgment is also an intervention
A judgment of learning, or JOL, asks a learner to predict future memory or performance. In a laboratory task the question might be, “How likely are you to remember this later?” In ordinary studying it appears in softer forms:
- Do I know this?
- How confident am I?
- Was that easy?
- Am I ready to move on?
- Which topic feels secure?
Researchers once often treated these reports mainly as windows into monitoring. A growing literature shows that asking for the report can itself change memory, attention or study behaviour. This is called reactivity.
Reactivity matters because a measuring instrument that changes the system must be interpreted differently. A thermometer that warms the liquid would not simply reveal temperature. In the same way, a confidence prompt can sometimes change the cognitive work being performed while it measures it.
The deep rule and the easy feature
Double, Tran and Goldwater studied relational category learning. Relational categories are defined by how elements relate, not merely by what they look like. “Larger than,” “chasing,” “inside,” “supports,” “causes” and “is the same proportion as” are relational ideas. Their surface forms can change while the relation remains.
That distinction is central to education. A mathematics problem may contain different numbers and objects while preserving the same ratio structure. A science example may use a different organism while preserving the same competitive relationship. A reading question may change the topic while preserving the same evidence-to-claim relation.
In the researchers’ first two experiments, participants could succeed using either a relational rule or visual features. Asking for JOLs impaired rule discovery but did not impair memorisation of the visual features. In a third experiment, where surface features could no longer provide a viable route and the relational rule was the only useful strategy, the negative effect was not observed.
The authors proposed a conservative strategy-shift account: when several strategies appear available, asking learners to evaluate their learning may push them toward the strategy whose success is easiest to see and reward immediately.
Why the easiest strategy can win
Imagine two routes through a task.
- Feature route: “Blue triangles usually belong to Group A.” It is concrete, fast and easy to check against recent examples.
- Relational route: “Group A contains cases where the smaller object is inside the larger object.” It requires comparison, abstraction and testing across changing surfaces.
If both routes happen to work during training, the feature route can produce rapid confidence. The relational route is harder to evaluate because the learner must test whether the relation survives across examples that look different.
A prompt such as “How well are you learning?” may therefore change the learner’s objective from discover what generalises to find evidence that I am succeeding now. That is a different optimisation problem.
Worked example: algebra
Illustrative case. A student practises equations such as 3x + 5 = 20, 7x + 2 = 30 and 4x + 9 = 25. The student notices that every worksheet example begins with a positive coefficient and has the variable term on the left.
The real relation is balance: whatever valid operation is performed to one side must preserve equality on the other. But the visible training pattern offers easier cues. If the student is repeatedly asked, “How confident are you with these?” confidence may rise because the worksheet format has become predictable.
Then the examination asks 18 = 3(2x − 1), or places variables on both sides. The surface cue disappears. The balance principle was never the thing being monitored.
A better monitoring question is not “How confident are you with equations?” It is: “What remains true when the equation looks different?”
Worked example: science
Illustrative case. A learner studies several food-web questions in which the producer is always drawn at the bottom and the predator at the top. The learner becomes fluent at reading the layout. A new diagram rotates the arrangement and uses unfamiliar organisms.
If learning was organised around picture position, confidence collapses. If learning was organised around the relation represented by the arrows, the learner can reconstruct the system. The same distinction appears throughout school: location versus function, wording versus logical role, notation versus invariant structure.
Recent evidence makes the picture more complicated, not less
The 2025 rule-discovery study is not evidence that every confidence rating damages learning. Other research finds positive, negative or null JOL reactivity depending on materials and conditions.
In June 2026, Ingendahl and Undorf reported three preregistered experiments using educationally relevant materials such as key-term definitions, country outlines and animal species. Immediate JOLs did not reliably improve later performance; in one experiment and an integrative analysis, performance was worse when difficulty was high. That result cautions against assuming that benefits found in simple paired-associate tasks automatically generalise to educational materials.
A January 2026 study by the same researchers found that instructing learners to use a strategy such as mental imagery could eliminate negative JOL reactivity for difficult unrelated word pairs. That suggests the monitoring prompt is not destiny. Strategy instruction can change what happens after the prompt.
Another 2026 study found JOL reactivity even when judgments were made covertly or with non-numerical response scales under conditions that engaged participants in the assessment. The broader point is that the cognitive act of assessing learning can matter, not only the act of typing a percentage.
Metacognition still matters
None of this overturns the strong educational case for teaching students to plan, monitor and evaluate their learning. The Education Endowment Foundation’s current guidance treats metacognition and self-regulation as valuable when strategies are explicitly taught, modelled and embedded in normal subject learning.
The distinction is between useful metacognition and unstructured self-rating. Monitoring should help the learner interrogate the learning goal. It should not simply generate more confidence numbers.
Monitor the mechanism, not only the feeling
Replace broad questions with questions tied to the capability you are trying to build.
- Instead of “Do I know this formula?” ask “When does this formula apply, and what condition would make it fail?”
- Instead of “Am I confident with this chapter?” ask “Can I classify a new problem whose surface looks different?”
- Instead of “Was that example easy?” ask “Which relation made the method valid?”
- Instead of “Can I recognise the answer?” ask “Can I produce it without the original cues?”
- Instead of “How much do I remember?” ask “What evidence would show that I can transfer this?”
The question should point attention at the structure you want the learner to own.
The competing-strategy diagnostic
When a learner performs well in training, ask whether more than one strategy could explain success.
- What obvious surface cue is available?
- What deeper relation is supposed to be learned?
- Could the learner succeed without understanding the relation?
- Can you build a new item where the surface cue points the wrong way?
- Does performance survive that item?
This is a powerful teaching move because it turns hidden strategy into observable evidence.
A surface-break test
After practice, deliberately break the familiar surface while preserving the deep relation. Change the numbers, orientation, names, diagram layout, wording or context. Then ask the learner to identify what stayed invariant.
If performance survives, the learner has stronger evidence of structural learning. If it collapses, do not label the student careless. The training may have rewarded an easier strategy.
When confidence ratings are still useful
Confidence can be valuable when it is compared with performance rather than treated as performance. A student can predict, attempt, receive feedback and then inspect calibration. The rating becomes one piece of evidence in a larger loop.
Useful uses include:
- identifying high-confidence errors;
- tracking whether calibration improves across delayed tests;
- deciding where to sample independent practice;
- noticing when confidence depends on familiar wording;
- comparing confidence before and after a surface-break transfer item.
The rating is most informative when the learner knows what capability the number refers to.
Delayed and independent performance check
Do not evaluate rule learning immediately beside the examples that revealed the rule. Wait, remove the training materials and present a structurally similar problem with changed surface features. Ask the learner to classify it, justify the classification and state the relation that controls the decision.
Then include a lure whose surface resembles the training examples but whose underlying relation is different. The pair of items distinguishes feature following from structural transfer much better than a confidence score alone.
For parents and tutors
Do not remove self-reflection from learning. Make it more specific. When a student says, “I know this,” ask what evidence supports the claim. Then change one irrelevant feature of the problem and one relevant structural feature on two separate examples. See whether the student knows which change matters.
When a child repeatedly chooses the obvious route, do not merely demand deeper thinking. Make the competition visible. Ask: “What clue are you using? Would that clue still work if I changed the picture? What rule would survive?”
What not to conclude
The laboratory findings do not prove that classroom reflection harms conceptual learning, nor that every JOL pushes every learner toward surface memorisation. The rule-discovery experiments used a controlled category-learning task. Effects of JOLs vary with material, difficulty, timing and strategy.
They do establish a more general methodological warning: self-report prompts are not guaranteed to be passive observations. When a learning system asks students to rate themselves constantly, the ratings may alter attention, goals and strategy. That possibility should be tested rather than assumed away.
The deeper studying principle
Good studying is not only about learning more. It is about learning the thing that will still matter after the surface changes.
Metacognition should help a learner find that thing. If monitoring instead rewards whatever feels easiest to judge, the instrument has taken control of the task. The solution is not less reflection. It is better-targeted reflection, tied to independent performance, structural transfer and the mechanism the learner is supposed to own.
Research basis and routes
Primary research: Double, K. S., Tran, D. & Goldwater, M. B. (2026 issue; published online 4 June 2025), “Judgments of learning impair rule-based discovery”, Memory & Cognition.
Current boundary evidence: Ingendahl, F. & Undorf, M. (2026), “Immediate judgments of learning do not improve performance for educationally relevant materials”; Ingendahl & Undorf (2026), “Instructed learning strategy use eliminates negative reactivity of immediate judgments of learning”; and the EEF Metacognition and Self-Regulated Learning guidance.
For calibration specifically, see How Studying Works | Delayed Judgments of Learning. Continue through the How Studying Works Numbered Series Reading Index or return to the How X Works Hub.