HSW-0240
A student performs poorly on a cognitive task. One adult says the score reveals weak ability. Another says the student was simply unmotivated. A third adds encouragement and feedback, the score rises, and everyone feels vindicated.
But the improvement does not automatically prove either story.
A test score is produced by a person performing a task under a set of conditions. The target capability matters. So do instructions, feedback, attention, strategy, fatigue, familiarity, stakes and other state variables. Some measures are more sensitive to those conditions than others. The scientific problem is to determine how much of the observed score reflects the construct we want and how much reflects the way the measurement was taken.
A 2026 Memory & Cognition study examined this problem across twelve cognitive tasks. Providing trial-, block- and task-level performance feedback plus normative comparison improved performance on all three attention-control tasks and two of three secondary-memory tasks. Yet primary-memory and fluid-intelligence tasks did not show the same task-level pattern. Self-reported motivation was only weakly related to performance, and the feedback manipulation did not significantly raise self-reported motivation. The authors therefore caution that the observed effects may have involved other consequences of feedback, such as monitoring or strategy selection.
This article calls the broader educational problem assessment-state sensitivity: the fact that some measured performances can move when the testing state or procedure changes, even though that does not make the underlying capability imaginary, nor justify explaining every low score as motivation.
The direct answer
A score is evidence about performance under specified conditions. It becomes evidence about a broader capability only to the extent that the task is valid for that claim and the result is stable enough across relevant conditions.
Feedback can change performance on some cognitive tasks without changing every task or reorganising the whole structure of measured abilities. Motivation can contribute to performance without accounting for most individual differences. A learner’s result should therefore be interpreted neither as a pure scan of fixed ability nor as a disposable mood reading.
The useful move is replication under controlled conditions: stabilise what should be stable, vary what you suspect matters, and see whether the same weakness returns.
Every assessment contains at least two systems
The first system is the capability being assessed: attention control, memory, reasoning, vocabulary, algebraic manipulation, reading comprehension or some other target.
The second is the performance environment that asks the learner to express that capability. It includes the task format, pacing, instructions, feedback, stakes, response mode and surrounding context.
A strong assessment makes the second system serve the first. A weak interpretation forgets the second system exists.
That is why changing a test condition can change a score. The crucial question is whether the change reveals a measurement artefact, changes how the learner deploys the same capability, teaches something during the assessment, or actually alters the target process being measured. Different mechanisms have different implications.
What the 2026 study actually manipulated
Stephen Campbell, Xavier Celaya, Alexis Torres, Gene Brewer and Matthew Robison recruited 674 participants for a battery of twelve cognitive tasks. The tasks were chosen to represent four domains: attention control, primary or short-term memory, secondary or long-term memory, and fluid intelligence.
Participants were randomly assigned to a feedback or no-feedback condition. In the feedback condition, people received performance information at multiple levels—depending on the task, trial-by-trial or block-by-block feedback as well as overall performance information and normative comparison with others.
The researchers intended this as a manipulation of motivation. Importantly, they also measured self-reported motivation before the experimental trials of each task.
The manipulation did not produce one uniform improvement. All three attention-control tasks showed significantly better performance in the feedback group. Two of the three secondary-memory tasks also improved. The primary-memory and fluid-intelligence tasks did not show significant task-level improvements from feedback.
That domain selectivity is central. If feedback simply released a large hidden reserve of general cognitive ability, we might expect a much more uniform shift. Instead, the measures differed in sensitivity.
The motivation story became more complicated
Participants rated how motivated they felt to perform well on each task. Those ratings were reliable across the session, but their relationships with performance were modest. At the task level, average absolute correlations were around .10 in the control condition and .14 in the feedback condition—less than two percent shared variance on average. At the latent level, the mean motivation rating correlated weakly with the cognitive factors, around .20 on average, corresponding to roughly four percent of variance.
Even more importantly, the feedback manipulation did not significantly change self-reported motivation on any task. The authors therefore explicitly note an alternative interpretation: where feedback improved performance, it may have operated through another mechanism such as metacognitive monitoring or strategy selection rather than through a successfully measured increase in motivation.
This is excellent evidence discipline. The study was designed around motivation, but the manipulation check did not support a simple “feedback increased motivation, which increased scores” chain. The behavioural effect remains real under the experimental conditions; the causal label becomes more cautious.
A score can be sensitive without being meaningless
Students and parents sometimes hear that test performance depends on context and jump to the conclusion that scores are arbitrary. That does not follow.
In the 2026 study, feedback affected some individual tasks and factor loadings, yet the underlying relationships among the broad cognitive factors were relatively stable across conditions. The authors describe this as encouraging from a measurement perspective: motivational factors, as captured in this study, did not appear to replace the intended cognitive constructs with a large source of systematic variance.
So the correct picture is layered. A measure can capture a real capability while still containing state-sensitive variance. Measurement is not an all-or-nothing choice between “pure ability” and “mere circumstances.”
Illustrative case: one unexpectedly low diagnostic
The following case is illustrative, not research data.
Daniel usually solves multi-step algebra accurately. On a short diagnostic, he performs badly. He misses several simple signs, leaves two items blank and appears slow.
There are many possible explanations. His algebra knowledge may have weakened. He may have misunderstood the response format. He may be fatigued. He may have allocated attention poorly. The low-stakes setting may have changed effort. He may be using a bad checking strategy. One unusual cluster of item types may have exposed a genuine gap.
A poor diagnostic response is not to select the explanation that best protects Daniel’s self-image. It is to design a second observation that discriminates among the explanations.
The tutor checks instructions, gives a short rest, uses a fresh but equivalent set, removes unnecessary time pressure and asks Daniel to show the first decision on each question. If the same sign-control failures return, the case for a real skill weakness strengthens. If performance normalises but collapses only under sustained timed work, the target has shifted toward performance control under load.
One score triggered a hypothesis. Repeated, controlled evidence identified the job.
Why feedback can change a test without teaching the tested content
Feedback during a cognitive task can alter several processes.
- It can increase awareness of errors.
- It can reveal whether a response strategy is working.
- It can change speed–accuracy trade-offs.
- It can sustain attention by making performance more salient.
- It can create social comparison through normative information.
- It can increase or decrease motivation depending on how it is received.
- It can teach the learner something about the task’s response structure.
That means “feedback improved the score” is a behavioural observation, not yet a complete mechanism.
This distinction matters when designing school diagnostics. If feedback is provided during a test intended to estimate unaided capability, the assessment may become part measurement and part intervention. That can be useful if the goal is dynamic diagnosis—seeing how the learner responds to help—but it changes the claim.
Measurement claims need matching evidence
The existing eduKateSG guide How Evidence-Centered Design Works begins assessment design with the claim we want to make, the evidence that claim requires and the task capable of producing it. HSW-0240 adds a state-sensitivity question: under what conditions should the claimed capability be expected to appear?
If the claim is “the learner can solve this independently under ordinary exam conditions,” then hints and correctness feedback should not be mixed into the decisive measurement. If the claim is “the learner can respond productively to feedback,” then feedback is part of the construct being investigated. If the claim is “the learner knows the method but loses control under sustained attention demand,” the test must contain enough sustained demand to expose that problem.
Assessment conditions are not background decoration. They help define the evidence.
Do not use motivation as a rescue explanation
“He could do it if he cared” is one of education’s most dangerous unfalsifiable statements. It can excuse a weak teaching diagnosis, dismiss a real knowledge gap or convert frustration into a character judgement.
The 2026 study gives no support for using motivation as a universal explanation of low cognitive performance. Self-reported motivation accounted for only modest variance, and the relation differed by task. Feedback effects themselves could not be cleanly attributed to measured motivation.
Motivation matters. But to claim it caused a particular failure, the diagnosis should produce discriminating evidence. Does performance change when incentives, feedback, task value or goal salience change while content demands remain comparable? Does the learner know the method when prompted but fail to deploy it? Does the same weakness appear on high-value tasks the learner clearly wants to solve?
Motivation should be investigated, not invoked.
Do not use “ability” as a verdict either
The opposite error is treating one observed score as a permanent personal property.
A valid test can provide important evidence about current capability and individual differences. But even good measures have error, state sensitivity and task-specific features. The appropriate confidence of an interpretation depends on reliability, validity, repeated evidence and the consequence of the decision being made.
This is especially important in tutoring, where a diagnostic is usually intended to find a repairable learning constraint rather than classify a person. A score should narrow the search. It should not end it.
The related HSW article HSW-0236 on the self-regulation judgment blind spot makes a similar point from another direction: visible outcomes should not be allowed to stand in for hidden learning processes without diagnostic evidence.
A four-pass diagnostic protocol
Pass 1 — Observe the failure without explaining it
Record what happened: response accuracy, omissions, latency, error type, sequence and task conditions. Avoid labels such as careless, weak, lazy or anxious unless independently supported.
Pass 2 — Stabilise the obvious conditions
Clarify instructions, remove accidental distractions, ensure the learner understands the response format and use a comparable task. This tests whether the original score survives a cleaner measurement environment.
Pass 3 — Change one suspected state variable
If attention is suspected, shorten the block or vary sustained demand. If feedback use is suspected, compare unaided performance with a feedback condition. If time pressure is suspected, compare timed and untimed equivalents. Change one thing where possible, not six.
Pass 4 — Return after delay under target conditions
The learner eventually has to perform under the conditions that matter. After any support or repair, use a fresh delayed test that restores the target constraints. That is where learning and state control meet.
Stable weakness versus state-sensitive weakness
A stable weakness appears repeatedly across well-matched tasks and ordinary variations in testing condition. It may point toward missing knowledge, an inefficient method or a genuinely underdeveloped skill.
A state-sensitive weakness changes strongly when a specific condition changes: sustained attention, feedback availability, time pressure, response mode, sleep, anxiety or other relevant states. That does not make it unreal. It means the performance system has a conditional failure.
In an examination, a state-sensitive weakness can be every bit as consequential as a content gap. If knowledge disappears reliably under the clock, exam readiness still requires repair. But the repair differs from simply reteaching the chapter.
Delayed and independent performance check
After feedback, encouragement or strategy coaching improves a task, remove the support and wait.
Then present a fresh task measuring the same capability. Keep the target conditions stable. Ask whether the improvement remains when correctness cues, normative comparison or tutor prompts are gone.
If performance remains better, the intervention may have changed strategy or learning in a durable way. If performance drops back immediately, the support may have been functioning as a temporary performance scaffold. Both findings are useful. They simply answer different questions.
Then vary the surface. A learner who improved only on the original test may have learned the test rather than the underlying capability.
For parents: replace character explanations with testable questions
When a child produces an unexpectedly low result, avoid beginning with “You didn’t try” or “You just aren’t good at this.” Neither statement identifies a repair.
Ask instead: Was this result typical? What kinds of errors occurred? Did the child understand the task? Does the same pattern recur after rest? Does it recur on a different format? Does feedback change the response? Does the improvement survive when feedback disappears?
These questions are slower than a label but much more useful. They also protect the child from learning the wrong causal story about themselves.
For tutors and teachers: decide whether feedback belongs inside the measurement
If the goal is baseline diagnosis, let the learner attempt enough work unaided to expose the current state. If the goal is to discover responsiveness to instruction, introduce help deliberately and record how much is needed. Do not blend the two phases invisibly.
A useful diagnostic record can separate:
- unaided performance;
- performance after a general prompt;
- performance after specific feedback;
- performance after reteaching;
- delayed independent return.
Now the assessment tells a story of assistance and transfer instead of collapsing everything into one mark.
Misconceptions
“Scores are mostly motivation.” Not supported by the 2026 study. Self-reported motivation showed weak relationships with performance on average and explained only modest variance at the latent level.
“Feedback raised scores, so it raised motivation.” The researchers intended feedback as a motivation manipulation, but self-reported motivation did not differ significantly between groups. Other feedback mechanisms remain plausible.
“If conditions change the score, the test is invalid.” Too strong. Some state sensitivity is compatible with a measure still capturing meaningful individual differences. Interpretation depends on the construct and use.
“A repeat test always solves the problem.” No. Practice effects, memory for items and additional learning can change repeated performance. Use alternate but comparable tasks where possible.
“A learner who improves with feedback has mastered the skill.” Improvement under support is evidence about responsiveness. Mastery requires a later independent check under target conditions.
Evidence and limits
The primary source is Campbell, Celaya, Torres, Brewer and Robison, “A combined experimental/individual differences examination of the influence of motivation on cognitive ability assessments”, published in Memory & Cognition on 7 February 2026. The study included 674 participants and twelve tasks spanning attention control, primary memory, secondary memory and fluid intelligence.
Participants in the experimental group received performance and normative feedback. All three attention-control tasks and two of three secondary-memory tasks showed significantly better task performance with feedback; primary-memory and fluid-intelligence tasks did not show significant task-level effects. Self-reported motivation correlated only weakly with performance, and the manipulation did not significantly increase motivation ratings.
The authors caution that feedback may have changed performance through processes other than motivation, including metacognitive monitoring or strategy selection. They also report that broad relationships among the cognitive constructs remained relatively stable, even though some task performance and factor loadings differed by condition.
Several limits matter. This was a single-session laboratory study using specific cognitive measures. Self-reported motivation may not capture every relevant motivational state. Feedback and normative comparison are composite manipulations, so the study does not isolate one active ingredient. The results should not be used to diagnose an individual student or to claim that school examination scores have the same sensitivity profile.
For assessment interpretation more broadly, How Concept Inventories Work and How Evidence-Centered Design Works provide the adjacent validity framework: the evidence must match the claim, and a response pattern should not be treated as a direct scan of the learner’s mind.
The return
A score is not the learner. It is not nothing either.
It is an observation produced by a capability meeting a task under conditions. Good diagnosis respects all three parts.
When a score moves after feedback, do not rush to say ability changed. Do not rush to say motivation was the whole problem. Ask what the feedback changed, whether the effect appears on comparable tasks, and whether the improvement remains when the support is removed.
When a low score repeats under stable conditions, take that evidence seriously. When it changes systematically with a condition, take that evidence seriously too.
The aim is not to make every result uncertain. It is to make uncertainty useful enough to locate the next test, the next repair and the next independent proof.
Continue through the How Studying Works Numbered Series Reading Index, or return to the How X Works Hub.