HSW-0252 · How Studying Works
A student learns about cognitive bias.
She learns to distrust the first answer that comes to mind. She practises base-rate problems. She sees why the gambler’s fallacy is seductive. She learns that availability can make vivid events feel more probable than they are. She becomes better at classic rational-thinking tasks.
Then she reads an argument claiming that a new study proves a sweeping conclusion.
The sample is narrow. The comparison is weak. The conclusion reaches beyond the evidence. One example is treated as if it settles the general case.
She accepts it anyway.
How can someone improve at rational-thinking exercises and still fail to evaluate an argument?
Because becoming better at one family of reasoning problems does not automatically create the knowledge, triggers, criteria and procedures needed for another family—even when the two performances are correlated and both look like “critical thinking.”
This is the rational-thinking transfer gap.
A 2026 pair of preregistered experiments found that brief rational-thinking training improved performance on general rational-thinking tasks, yet those improvements did not transfer to a separate scientific argument-evaluation task. The result does not show that rational-thinking training is useless. It shows something more educationally valuable: training effects must be demonstrated at the performance you actually care about.
This article owns the narrow study question of why improvement on cognitive-bias and rational-thinking tasks may fail to transfer to argument evaluation. The broader canonical owner remains How Transfer of Learning Works. Cross-Domain Structural Priming retains the narrower question of whether processing one abstract structure can automatically prime another domain. Inert Knowledge retains the general case in which useful knowledge exists but is not activated when needed.
The 50-Second Read
- Related skills are not the same skill. Rational-thinking performance and argument evaluation can correlate while relying on partially different knowledge and operations.
- Correlation does not guarantee intervention transfer. Two abilities can travel together across people even when training one does not cause improvement in the other.
- Two preregistered 2026 experiments found near improvement without the hoped-for transfer. Rational-thinking training improved rational-thinking tasks, but not scientific argument evaluation relative to active controls.
- Transfer requires a usable bridge. Learners need cues that trigger the relevant knowledge, criteria for judging the new task and practice applying those criteria.
- Critical thinking is not one generic muscle. Bias checking, evidence evaluation, causal reasoning, source judgment and argument analysis overlap but are not interchangeable.
- Practice effects can imitate learning. Both treatment and control groups improved on argument evaluation, so repeated exposure itself may explain some gains.
- The practical rule: train the target epistemic act, then test it later on new content without announcing which reasoning tool to use.
1. The Temptation to Call Everything “Critical Thinking”
Education often compresses many different performances into one attractive label:
critical thinking.
But consider what that label may be asked to cover:
- resisting a tempting intuitive answer;
- using base rates correctly;
- detecting a contradiction;
- distinguishing correlation from causation;
- evaluating whether evidence supports a claim;
- spotting a false dichotomy;
- recognising circular reasoning;
- judging whether an example generalises;
- checking whether a source is trustworthy;
- deciding what evidence would change a conclusion.
These performances are related. They can share habits of slowing down, checking evidence and resisting easy answers.
They are still not identical.
A learner can become excellent at one and remain weak at another because the cue, representation, domain knowledge and decision rule differ.
2. What the 2026 Experiments Tested
A 2026 study in Learning and Instruction tested a straightforward educational hope: if people receive training in rational thinking, will that improvement transfer to evaluating scientific arguments?
The researchers ran two preregistered experiments with university students. The first included 149 participants and the second 155. Participants were assigned either to rational-thinking training or to an active control activity.
The rational-thinking training addressed ideas associated with dual-process accounts of cognition and common reasoning biases. Participants practised tasks involving intuitive-versus-reflective responding and phenomena such as availability, gambler’s-fallacy reasoning and base-rate use.
Rational-thinking performance was assessed with tasks including cognitive-reflection-style problems, belief-bias resistance, ratio-bias problems and disjunctive reasoning. Argument evaluation was measured separately using arguments containing weaknesses such as circularity, false dichotomies, inappropriate examples, overgeneralisation and contradiction.
The treatment improved rational-thinking performance in both experiments.
But it did not produce a corresponding treatment advantage in argument evaluation.
Both treatment and control groups improved on argument evaluation, suggesting that repeated exposure, practice with the assessment or other non-specific factors may explain that change.
Read the source: Training rational thinking: Improvements without transfer to argument evaluation, 2026.
3. The Result Is More Interesting Than “Training Failed”
The rational-thinking training did not fail at its immediate job.
Participants improved on the rational-thinking measures.
The failure was at the bridge.
Improvement did not automatically travel from the trained family of reasoning tasks into scientific argument evaluation.
That distinction matters because educational programmes are often sold or defended with a chain like this:
we trained a reasoning skill → reasoning is important everywhere → therefore the training improves reasoning everywhere.
The middle statement may be true while the final inference is not demonstrated.
4. Correlation Is Not Transfer
The 2026 study found that rational-thinking skill and argument-evaluation performance were related.
That sounds, at first, like transfer should follow.
But cross-sectional association and intervention transfer answer different questions.
Correlation asks:
Do people who perform better on A also tend to perform better on B?
Transfer asks:
If we deliberately improve A, does B improve because A improved?
Those are not interchangeable.
A third factor can support both A and B. People with more education, stronger domain knowledge, better reading comprehension or greater willingness to reflect may perform well on both. Training one measured manifestation need not alter the other.
A trait relationship tells us what travels together across people. A transfer experiment tells us what travels after training.
5. The Missing Trigger Problem
A learner can know a reasoning principle and fail to activate it.
During training, the task itself may announce the relevant move:
- this is a base-rate problem;
- this is a cognitive-reflection item;
- this is a probability trap;
- this is a bias exercise.
In real argument evaluation, nobody announces the category.
The learner must first notice that something deserves scrutiny, classify the weakness and retrieve the appropriate criterion.
That trigger can be the missing transfer component.
6. The Missing Criteria Problem
Being generally reflective does not tell you what makes a scientific argument good.
Argument evaluation requires standards.
- Does the conclusion actually follow?
- Is the evidence relevant?
- Is the sample appropriate?
- Is a causal claim justified?
- Are alternatives considered?
- Is the generalisation wider than the data?
- Is the argument circular?
- Does the evidence support the strength of the wording?
Without these criteria, “think more carefully” is under-specified.
7. The Missing Knowledge Problem
Argument quality often depends on domain knowledge.
Suppose a claim says a drug “caused” improvement because treated patients improved after taking it.
Evaluating that statement well may require knowledge of control groups, regression to the mean, placebo effects, measurement, random assignment and alternative causal explanations.
A general lesson about cognitive bias does not automatically supply that methodological knowledge.
Critical thinking cannot operate independently of what there is to think with.
8. The Missing Procedure Problem
Even when a learner knows the criteria, they may not have a routine for applying them.
A useful argument-evaluation procedure might be:
- state the exact claim;
- identify the evidence offered;
- state the inferential bridge;
- test whether the bridge is warranted;
- search for alternatives;
- calibrate the conclusion to the evidence;
- state what additional evidence would change the judgment.
Without a procedure, abstract caution may never become observable performance.
9. Mathematics Example: Reflection Is Not Proof Evaluation
A student can become better at resisting a tempting numerical answer and still accept an invalid proof.
Why?
Proof evaluation requires mathematics-specific criteria:
- was every transformation valid?
- were domain restrictions preserved?
- was a special case treated as a general proof?
- was a reversible implication silently assumed?
- did the argument establish necessity, sufficiency or both?
General reflectiveness may increase the chance of pausing. It does not supply the proof standard automatically.
10. English and GP Example: Spotting a Bias Is Not Evaluating Evidence
A student may correctly define confirmation bias yet write:
“Social media causes political polarisation because people often encounter extreme opinions online.”
The problem is not necessarily that the student forgot confirmation bias.
The student may lack a disciplined argument-evaluation routine:
- What exactly is the causal claim?
- What evidence distinguishes cause from selection?
- Could polarised people choose more extreme content?
- How strong is the effect?
- For which populations and platforms?
Knowing the vocabulary of bias is not the same as performing evidence evaluation.
11. Science Example: Cognitive Reflection Is Not Methodological Judgment
A learner may solve a classic reflection problem by suppressing the intuitive answer.
Then the learner reads a study summary and fails to notice that the outcome was measured only once, the groups were not comparable or the conclusion extends beyond the sample.
Scientific reasoning needs domain-relevant methodological knowledge in addition to the disposition to think again.
12. Rational-Thinking Transfer vs Cross-Domain Structural Priming
Cross-Domain Structural Priming asks whether processing one structural pattern automatically biases subsequent processing of an analogous structure in another domain.
The rational-thinking transfer gap is broader and instructional. It asks whether deliberate training on one reasoning family produces measurable improvement on a different epistemic performance.
Both warn against automatic transfer, but they test different mechanisms.
13. Rational-Thinking Transfer vs Inert Knowledge
Inert Knowledge explains how relevant knowledge can remain unused because the learner does not recognise when to retrieve it.
That may contribute to the transfer gap, but it is not the only explanation. The learner may also lack the target criteria, the procedural routine or the required content knowledge.
Transfer diagnosis should therefore avoid treating every failure as a cue problem.
14. Why Active Controls Matter
If a training group improves from pre-test to post-test, it is tempting to credit the training.
But improvement can come from:
- seeing the test before;
- becoming familiar with the interface;
- learning what the questions demand;
- simple regression toward the mean;
- greater motivation at post-test;
- general practice unrelated to the treatment’s distinctive mechanism.
An active control helps separate the unique contribution of the treatment from participation and practice effects.
In the 2026 experiments, both groups improved in argument evaluation. That pattern is precisely why the treatment comparison matters more than the before-and-after story alone.
15. Measurement Matters Too
A null transfer result can mean several things.
- transfer truly did not occur;
- transfer was too small for the study to detect;
- the target measure was not sensitive enough;
- measurement reliability was limited;
- training dosage was insufficient;
- the trained and target tasks shared less than theory assumed.
The argument-evaluation measure in the study included several types of argument weakness, and some subcomponents had weaker reliability than ideal. That limits how finely the null result can be interpreted.
The correct conclusion is not “rational-thinking training can never help argument evaluation.” It is that these brief interventions did not demonstrate such transfer under the tested conditions.
16. The Transfer Bridge Audit
When two skills seem related, inspect the bridge rather than assuming it.
| Bridge component | Question |
|---|---|
| Trigger | Will the learner recognise when the old reasoning tool is relevant? |
| Representation | Can the new problem be expressed in a form compatible with the old knowledge? |
| Criteria | Does the learner know what counts as a good answer in the target domain? |
| Knowledge | Does the learner have the content needed to evaluate the target claim? |
| Procedure | Is there an executable sequence for applying the reasoning? |
| Discrimination | Can the learner tell when the old tool does not apply? |
| Retrieval | Can the method be accessed without a teacher naming it? |
| Verification | Can transfer be demonstrated on new delayed tasks? |
17. The Target-to-Target Training Protocol
If the real goal is argument evaluation, train argument evaluation.
- Name the target performance. “Evaluate whether the evidence supports the claim” is better than “think critically.”
- Teach the criteria. Relevance, validity, alternatives, scope, source quality and conclusion strength.
- Model the procedure. Claim → evidence → bridge → alternative → calibrated conclusion.
- Use contrast cases. Strong and weak arguments should differ on one decisive feature.
- Fade the labels. Stop announcing which fallacy or bias is present.
- Mix argument types. Some examples should contain no flaw.
- Change subject matter. Move from health to economics to education to science.
- Delay the test. Return after the training context has faded.
- Require justification. A correct verdict without a reason may hide guessing.
- Test transfer again. Use new claims with unfamiliar surface features.
18. Train the Trigger, Not Only the Answer
Transfer fails when the learner knows what to do only after the teacher asks the right question.
Train triggers such as:
- a strong causal verb;
- a sweeping universal claim;
- an emotionally vivid example;
- a tiny or selective sample;
- a comparison without a baseline;
- a binary choice in a multi-option situation;
- evidence that repeats the conclusion rather than supports it.
The trigger should initiate evaluation before the learner is told that “this is a critical-thinking question.”
19. Train Disconfirmation
One of the strongest bridges from abstract rationality to real evaluation is the ability to ask what would make a conclusion weaker.
The existing eduKateSG guide Decide What Evidence Would Make You Change Your Answer owns that broad examination-performance procedure.
In argument evaluation, the same discipline becomes:
- What observation would contradict this explanation?
- What alternative mechanism fits the same evidence?
- What missing comparison would matter?
- What population would challenge the generalisation?
- What result would make the conclusion too strong?
20. The Near-to-Far Reasoning Ladder
- Practise one named bias with immediate feedback.
- Mix several named biases.
- Remove the bias labels.
- Include sound arguments that should be accepted.
- Move from puzzle-like items to short real arguments.
- Change topic while preserving the reasoning issue.
- Add irrelevant but plausible detail.
- Use longer authentic sources.
- Test later without announcing that argument evaluation is the target.
Each rung removes one support that may have been carrying performance.
21. The Delayed Independent Check
A reasoning lesson is not finished when students can explain the bias at the end of class.
Return several days later with a new argument embedded inside ordinary subject work.
Do not label the flaw.
Ask the learner to:
- state the claim;
- identify the evidence;
- judge the inferential bridge;
- name the strongest alternative explanation;
- calibrate the conclusion;
- state what further evidence would matter.
If the learner can do this without the original training cues, the reasoning has begun to travel.
22. Parent and Tutor Guide: Do Not Ask Only “Do You Know the Bias?”
A learner may be able to define confirmation bias, base-rate neglect or the gambler’s fallacy perfectly.
That is useful vocabulary, not yet transferable judgment.
Ask instead:
- Can you find a weak inference when nobody names the bias?
- Can you distinguish a flawed argument from a strong one?
- Can you explain why the evidence is insufficient?
- Can you identify a missing comparison?
- Can you say what evidence would change your judgment?
- Can you repeat the performance on a different topic next week?
That sequence converts knowledge-about-reasoning into reasoning-in-use.
23. For Teachers: Stop Treating Transfer as a Compliment
“This activity develops critical thinking” is not a description of evidence.
Replace it with a testable claim:
After this training, students should become better at evaluating whether evidence justifies a conclusion in unfamiliar arguments.
Now the curriculum has a target, a transfer distance and an assessment.
If improvement appears only on the training puzzles, say so. If it survives to new arguments, say that instead.
24. The Systems Route: General Capability Claims Need Specific Acceptance Tests
Organisations make the same mistake as classrooms.
A programme teaches “decision quality,” “leadership,” “systems thinking” or “analytical reasoning,” then assumes the capability will appear wherever needed.
A stronger system defines acceptance tests:
- Which decisions should improve?
- Under what uncertainty?
- With which information?
- Against what standard?
- How long after training?
- Without which prompts?
Broad labels require narrow tests.
25. The Resource-Allocation Route: Spend Training Time Where the Final Error Occurs
Transfer is expensive.
If the examination, profession or life task requires evaluating arguments, then some training time should be spent evaluating arguments rather than assuming adjacent cognitive exercises will carry the load.
This does not mean every practice item must look like the final assessment.
It means the programme must eventually include the target decision under realistic cue conditions.
26. Competing Explanations for the Null Transfer
A rigorous interpretation keeps several explanations alive.
- Task specificity: the trained reasoning operations were not sufficiently shared with argument evaluation.
- Cue failure: useful principles were learned but not spontaneously triggered.
- Knowledge gap: argument evaluation depended on epistemic or methodological knowledge not taught in the intervention.
- Dosage: the brief training may have been too short to create broad transfer.
- Measurement: the target measure may have had limited sensitivity or reliability for some argument types.
- Practice effects: repeated argument testing may have raised both groups, shrinking the detectable treatment difference.
The experiment cannot uniquely select among all of these mechanisms. That is why the public conclusion should remain narrower than the educational imagination.
27. What Not to Conclude
- Do not conclude that rational-thinking training is ineffective; it improved rational-thinking tasks.
- Do not conclude that critical thinking cannot transfer under any conditions.
- Do not infer that argument evaluation and rational thinking are unrelated; the study found associations between them.
- Do not interpret a null treatment difference as proof of exactly zero transfer.
- Do not call repeated-test improvement proof that the intervention worked.
- Do not assume one short university-student intervention generalises to children, schools or long-term programmes.
- Do not replace subject knowledge with generic reasoning drills.
28. Evidence Boundary
The two experiments were preregistered and used active controls, strengthening causal interpretation relative to simple pre/post designs. The participants were university students, the training was brief and computer-based, and transfer was tested to one important but limited epistemic outcome: scientific argument evaluation.
Some subscales of the argument measure had weaker reliability, and the study does not test every form of critical thinking, every instructional design or long-term transfer. The strongest supported claim is therefore specific: improvement on the trained rational-thinking tasks did not produce a detectable treatment-specific improvement in the tested argument-evaluation performance.
29. The Improvement Rule: Measure the Destination
If a training programme claims transfer, measure the destination.
Not just the exercise.
Not just the vocabulary.
Not just confidence.
Not just a correlated trait.
Measure the performance the learner will eventually need when the original training cues are gone.
30. Return: Better Thinking Is Built at the Point Where Thinking Must Work
There is no reason to abandon rational-thinking instruction.
There is every reason to become more precise about what it buys.
A learner can become better at resisting one family of cognitive traps and still need direct training to evaluate evidence, arguments, sources and causal claims.
Teach the principle. Teach the trigger. Teach the criteria. Practise the target performance. Remove the labels. Wait. Then ask the learner to recognise and execute the reasoning independently.
That is the bridge from knowing about rationality to using rationality where it matters.
Continue through How Transfer of Learning Works, Cross-Domain Structural Priming, Inert Knowledge, the How Studying Works Numbered Series Reading Index and the How X Works Hub.