Why the Same Wrong Answer Can Have Different Causes
PMRI-003 — Primary Mathematics Research Institute Series
eduKateSG Learning Sciences & Education Research Programme
Framework: Weak-Link Atlas WLA v1.0
Status: Diagnostic Taxonomy / Research Framework
Research Domain: Primary Mathematics / Error Analysis / Diagnostic Education
Last Reviewed: August 2026
Research Classification
Article Type: Diagnostic taxonomy + evidence synthesis + proposed validation framework
Primary Research Question:
Can recurring Primary Mathematics failure patterns be classified into a small enough set of intervention-useful categories without pretending that one visible error has one unique cause?
Secondary Research Question:
Does classifying the mechanism behind an error lead to better intervention decisions than simply classifying the question as right or wrong?
Research Upgrade
The newest research review changes one part of the eduKateSG model substantially.
We should not write:
Error = Weak Link
Instead:
Error = Observation
and:
Weak Link = Hypothesis about the mechanism that generated the observation
A 2025 ZDM scoping review of recent research on Mathematics classroom errors distinguishes descriptive/diagnostic, intervention-oriented and inquiry-oriented uses of student errors. It emphasises that errors can reveal gaps and misconceptions, but also that mathematical errors, misunderstandings and mistakes are not interchangeable phenomena and can be influenced by cognitive, instructional and affective conditions. (Springer)
Therefore PMRI-003 upgrades the earlier taxonomy into three layers.
The Three-Layer Error Architecture
Layer 1 — Observed Failure
What actually happened?
Examples:
- wrong answer;
- unusually slow answer;
- wrong method selected;
- unit omitted;
- question left blank;
- excessive prompting required;
- answer changed repeatedly;
- correct answer obtained only after an example was shown.
This is observation.
Layer 2 — Proposed Weak-Link Mechanism
What part of the Mathematics system most plausibly generated the failure?
Candidate categories:
- Missing Node
- Broken Edge
- Weak Link
- Wrong Edge
- Routing
- Translation
- Transfer
- Calibration
- Regulation
These are hypotheses.
Layer 3 — Contributing Conditions
What conditions might strengthen or weaken the mechanism?
Examples:
- arithmetic fluency;
- task complexity;
- language demand;
- prior knowledge;
- prompt level;
- time pressure;
- confidence;
- attention;
- unfamiliar representation.
This distinction prevents the taxonomy from becoming simplistic.
Quick Read
Suppose three children all answer the same Mathematics question incorrectly.
Child A
Does not understand the concept.
Child B
Understands the concept but retrieves an old fact incorrectly.
Child C
Understands everything but chooses the wrong method.
Same:
wrong answer
Different:
system failure
Therefore:
Wrong Answer
↓
Find First Divergence
↓
Generate Candidate Mechanisms
↓
Collect More Evidence
↓
Intervene
↓
Retest
The purpose of the Weak-Link Atlas is not to give every mistake a fancy name.
Its purpose is to help answer:
Where should repair begin?
Part 1 of 3
Error Is Evidence, Not Diagnosis
The Binary Problem
Most assessments necessarily begin with:
correct
or:
incorrect
That distinction is useful.
But it loses information.
Consider:
Response A
Correct answer.
Correct reasoning.
Independent.
Response B
Correct answer.
Incorrect reasoning.
Lucky calculation.
Response C
Wrong answer.
Correct concept.
Arithmetic slip.
Response D
Wrong answer.
Wrong conceptual structure.
If all we store is:
1 / 0
we collapse four different learning states into two categories.
A research-oriented teaching system should reopen that compression where intervention decisions matter.
The First-Divergence Principle
Instead of asking:
Where is the final wrong answer?
ask:
Where did the solution first stop matching a valid mathematical path?
This may occur in:
- understanding the language;
- identifying the quantities;
- choosing a representation;
- selecting a method;
- retrieving a prerequisite;
- sequencing the steps;
- executing arithmetic;
- checking.
The first divergence often gives a better starting point for diagnosis.
Not always.
But often enough to investigate.
Recent Research Supports Deeper Error Analysis
The 2025 ZDM review describes student errors as potential indicators of mathematical thinking rather than merely objects to be corrected. It also notes that cognitive diagnostic approaches seek to identify persistent misconceptions so instruction can respond more precisely. (Springer)
This aligns closely with the direction of the eduKateSG programme.
But we add a boundary:
The error itself does not prove the mechanism.
We need additional observations.
Taxonomy Rule 1
Never Infer Cause From One Error Alone When Multiple Causes Are Plausible
Suppose:
7 × 8 = 54
Possible explanations include:
- fact not known;
- fact known but retrieved incorrectly;
- temporary attention failure;
- transcription error;
- confusion between nearby multiplication facts.
One response is insufficient for strong diagnosis.
Repeated evidence changes that.
Taxonomy Rule 2
Repetition Raises Diagnostic Value
Suppose across several sessions:
7 × 8
is slow.
6 × 8
is slow.
9 × 7
is slow.
Multiplication facts are repeatedly reconstructed.
Now the hypothesis:
Weak multiplication retrieval
receives more support.
The pattern matters.
Taxonomy Rule 3
Intervention Response Is Part of Diagnosis
Suppose we strengthen multiplication retrieval.
Then later fraction work becomes substantially faster and more accurate.
That supports the original weak-link hypothesis.
If multiplication becomes fluent but the fraction difficulty remains unchanged:
revise the diagnosis.
The intervention itself produces diagnostic evidence.
The Weak-Link Atlas
Now we can define the nine candidate mechanisms.
1. Missing Node
Definition
A mathematical concept, fact, representation or relationship required by the present task is not sufficiently available.
Example:
A learner is asked to compare equivalent fractions but has not established fraction equivalence.
The problem is not:
weak exam technique.
The relevant mathematical node is missing.
Missing Node Indicators
Possible signals:
- cannot explain concept even with simpler numbers;
- does not recognise the idea when represented differently;
- extensive prompting does not reconstruct it;
- errors remain systematic across close variants.
Missing Node Test
Simplify the surrounding task.
If the learner still cannot access the target concept:
Missing Node becomes more plausible.
Falsification Condition
If the student can explain, represent and apply the concept independently in simpler and varied contexts:
the node probably exists.
Search elsewhere.
2. Broken Edge
Definition
Two mathematical ideas exist individually but the learner does not reliably connect them.
Example:
The child knows:
multiplication
and:
division
but does not use known multiplication facts to reason about division.
Both nodes exist.
The edge is weak or absent.
Other Examples
fraction ↔ decimal
decimal ↔ percentage
multiplication ↔ area
factor relationships ↔ fraction simplification
Broken Edge Test
Present the two ideas separately.
Then present a task requiring their connection.
If separate performance is strong but integrated performance repeatedly fails:
Broken Edge becomes plausible.
3. Weak Link
Definition
Knowledge or a connection exists but is too unstable, slow or unreliable for current task demands.
This is different from Missing Node.
The learner can sometimes use the Mathematics.
But reliability is inadequate.
Example
The student knows multiplication facts but retrieves them inconsistently.
Or:
The learner understands equivalent fractions but loses control under multi-step load.
Weak-Link Test
Reduce time pressure and task complexity.
If performance improves substantially:
the capability may exist but lack sufficient stability or dispatchability.
4. Wrong Edge
Definition
The learner has connected a cue, representation or problem type to an inappropriate mathematical relationship.
Example:
“More” always means add.
The learner has built a shortcut.
But the shortcut is mathematically unreliable.
Wrong Edge Example
Question:
Ben has 5 more marbles than Ali.
The student automatically adds the two visible numbers because:
more → add
The problem is not missing addition.
The problem is an incorrect relationship between language and operation.
Wrong Edge Test
Construct contrasting examples using the same surface cue but different mathematical structures.
If the student repeatedly follows the cue rather than the relationship:
Wrong Edge becomes plausible.
5. Routing Gap
Definition
The learner possesses multiple potentially relevant mathematical strategies but cannot reliably select the appropriate one.
This becomes increasingly important as the curriculum expands.
Blocked Practice Can Hide Routing Weakness
Worksheet heading:
Fractions
already tells the student where to route.
Mixed paper:
no such support.
Now the learner must determine the mathematical family independently.
Routing Test
Mix problems from several previously learnt families.
Do not announce the topic.
Measure:
- method selection;
- latency;
- changes of route;
- need for prompts.
If knowledge is strong in blocked conditions but method selection collapses in mixed conditions:
Routing Gap becomes plausible.
6. Translation Gap
Definition
The learner has relevant Mathematics but cannot reliably transform language, notation or context into the required mathematical relationship.
Example
Direct question:
34 − 17
correct.
Word problem expressing the same relation:
wrong.
This is not enough to prove a Translation Gap.
But it creates a hypothesis.
Translation Test
Hold the mathematical structure constant.
Change:
symbolic
→
verbal
→
diagrammatic
If mathematical performance changes strongly with representational language:
translation becomes a likely factor.
Language and Mathematics Are Not Completely Separate
Recent Mathematics research continues to distinguish surface translation from deeper conceptual meaning-making. A 2025 study of adaptive task selection describes the importance of moving beyond superficial correspondences toward underlying mathematical structures in multiplication. (Springer)
This matters for eduKateSG because:
a language-looking error can sometimes be a mathematical-structure error.
We should test both possibilities.
7. Transfer Gap
Definition
Learning works within the training conditions but fails when the problem changes sufficiently.
Examples:
- same method, new wording;
- same concept, different diagram;
- familiar procedure, new context;
- concept appears alongside another topic.
Transfer Gap Test
Change one surface dimension at a time.
Then progressively combine changes.
If performance is strongly tied to the trained format:
transfer remains weak.
Transfer Is Not Optional
If the learner can solve only:
the question that looks like the example
then knowledge has limited portability.
Primary Mathematics eventually requires students to operate across mixed and unfamiliar problems.
So transfer belongs inside the weak-link model.
8. Calibration Gap
Definition
The learner cannot reliably judge whether a result is plausible or whether a solution path is malfunctioning.
Examples
A child accepts:
pencil length = 400 metres.
Or:
20% of a positive quantity is greater than 200% of the same quantity.
The numerical output exists.
Internal error detection does not activate.
Calibration Test
Introduce a deliberate impossible or highly implausible result.
Ask:
Does anything seem wrong?
A learner with stronger calibration should increasingly use:
- magnitude;
- units;
- context;
- relationships;
to challenge output.
9. Regulation Gap
Definition
Mathematical capability exists, but execution becomes unreliable because the learner cannot adequately regulate attention, time, checking, prompting dependence or recovery.
Examples
- repeatedly rushing the opening questions;
- spending excessive time on one blocked problem;
- abandoning checking completely;
- needing immediate adult reassurance after uncertainty;
- allowing one difficult question to affect several subsequent questions.
Regulation Is Not a Personality Label
The relevant statement is not:
“Careless child.”
It is:
Under these conditions, this execution pattern repeatedly appears.
That makes the problem observable and potentially trainable.
What About Arithmetic Mistakes?
The updated model deliberately does not create a tenth category called:
Execution Error
for every calculation slip.
Why?
Because:
Execution Error
is first an observed event.
The underlying mechanism may be:
- Weak Link;
- Missing Node;
- Regulation;
- or simple noise.
We should not mistake the response phenotype for the causal category.
This is one of the major upgrades from the new research scan.
Part 2 of 3
The Weak-Link Diagnostic Protocol
Step 1 — Record the Error Without Interpretation
Example:
Student wrote 7 × 8 = 54 during a three-step fraction question.
Do not immediately write:
weak multiplication.
Step 2 — Locate First Divergence
Was the strategy correct before the multiplication error?
If yes:
conceptual routing may be intact.
The first divergence is arithmetic.
Step 3 — Generate Competing Hypotheses
Possible explanations:
H1 — multiplication node missing
H2 — multiplication retrieval weak
H3 — regulation under load
H4 — isolated mistake
Keep several possibilities alive.
Step 4 — Design a Discriminating Probe
Ask several multiplication facts.
Untimed.
Then mixed.
Then later.
Observe.
The probe should separate competing explanations.
Step 5 — Assign Diagnostic Confidence
Low
One ambiguous event.
Moderate
Repeated pattern.
Higher
Multiple observations converge and intervention response supports the mechanism.
Never confuse confidence with certainty.
Step 6 — Intervene at the Proposed Weak Link
If retrieval is the hypothesis:
train retrieval.
If concept is missing:
rebuild concept.
If routing is weak:
train discrimination.
Step 7 — Look for Downstream Change
The purpose of a weak-link hypothesis is not only to fix the isolated probe.
Ask:
Did the original task become easier?
If not:
the proposed weak link may not have been the real bottleneck.
High-Leverage Weak Links
Not every weakness deserves equal priority.
A weakness becomes high leverage when it constrains many downstream tasks.
Example:
weak multiplication structure
may interfere with:
- division;
- fractions;
- percentage;
- area;
- volume.
Repairing it may yield broader benefit.
But this is still a hypothesis until spillover is observed.
Important Research Upgrade: Spillover Is Not Guaranteed
A 2025 Grade 5 foundation-intervention study provides a useful warning. Students receiving a conceptually focused intervention improved on the targeted basic concepts, but the study did not find automatic transfer to basic skills. (ResearchGate)
That matters greatly for eduKateSG.
We should not write:
Repair foundational concept A, therefore all downstream performance will automatically improve.
Instead:
Repair A, then test whether B changes.
This converts the weak-link framework from rhetoric into research.
Weak-Link Propagation Hypothesis
The refined hypothesis becomes:
Some prerequisite weaknesses constrain later Mathematics, but repairing the prerequisite does not guarantee automatic transfer into every dependent skill. Downstream change must be measured.
This is much stronger.
False Positive Weak Links
A diagnosis can be wrong.
Example:
The child fails a ratio question.
Tutor assumes:
fractions are weak.
Fractions are then retaught extensively.
Ratio remains weak.
The problem may have been:
- translation;
- routing;
- reference quantity;
- or ratio concept itself.
The framework must therefore record false positives.
False Negative Weak Links
The opposite can occur.
A child appears fluent during direct practice.
Tutor concludes:
multiplication is stable.
But mixed multi-step work reveals severe retrieval friction.
The weakness existed but the original probe failed to detect it.
This is a false negative.
Research should examine both.
Error Genealogy
One powerful future direction is to trace errors across time.
For example:
P3 multiplication retrieval weakness
↓
P4 factor instability
↓
P5 fraction inefficiency
↓
P6 percentage/ratio overload
This is a proposed genealogy.
It should not be inferred retrospectively without evidence.
But longitudinal records may eventually allow us to test whether particular error families propagate in predictable ways.
That is a major research opportunity for eduKateSG.
Part 3 of 3
From Taxonomy to Research Instrument
The Weak-Link Atlas should eventually be validated.
Questions include:
Can Tutors Distinguish Categories Reliably?
If two trained tutors inspect the same work:
do they identify similar candidate mechanisms?
Do Categories Predict Intervention Response?
If a problem is classified as Routing:
does routing-focused intervention outperform generic practice?
Are Some Categories Redundant?
Perhaps:
Broken Edge
and:
Wrong Edge
cannot be reliably separated in some contexts.
Then revise the model.
Are Categories Topic-Specific?
Perhaps Translation behaves differently in geometry and fractions.
Investigate.
Design-Based Research Fits the eduKateSG Programme
A 2025 Educational Studies in Mathematics paper describes design-based research as an iterative process in which interventions and local theories are developed and revised through cycles of empirical investigation and retrospective analysis. (Springer)
That is a particularly good methodological fit for eduKateSG.
The Weak-Link Atlas should therefore be treated as:
a local theory under iterative development
not:
universal final truth.
WLA Versioning
WLA v1.0
Nine candidate weak-link classes.
Future WLA v1.1
May alter:
- definitions;
- indicators;
- category boundaries;
- probe methods.
WLA v2.0
Should require stronger empirical validation before claiming broader diagnostic use.
Version history should record:
What changed?
Why?
What evidence required the change?
Parent-Facing Translation
The research taxonomy should not force parents to learn technical language.
Instead of:
“Your child has a Broken-Edge pathology.”
say:
“The child understands multiplication and division separately, but does not yet use them as connected inverse relationships. We are working on that connection.”
The technical framework stays behind the educational explanation.
Student-Facing Translation
Instead of:
“You are weak at fractions.”
say:
“You understand the fraction itself. The part we are strengthening is deciding when to use that idea inside a mixed question.”
The taxonomy should reduce identity labels.
Not create new ones.
The Weak-Link Atlas Runtime
Observe
↓
Describe Without Interpretation
↓
Find First Divergence
↓
Generate Competing Mechanisms
↓
Probe
↓
Triangulate
↓
Assign Confidence
↓
Select Highest-Leverage Candidate
↓
Intervene
↓
Retest Target
↓
Test Downstream Effect
↓
Delay
↓
Transfer
↓
Update Diagnosis
↓
Record False Positives / False Negatives
↓
Revise Atlas
Evidence Boundary
Current Mathematics education research supports careful analysis of errors and misconceptions as potentially useful evidence for instructional decision-making, while also showing that errors arise within broader student–teacher–content interactions and should not be reduced to a simple right/wrong binary. (Springer)
Research on learning trajectories and diagnostic assessment likewise supports using evidence-based cognitive models to identify relevant prior knowledge and possible misconceptions. (Springer)
What is not yet established is that eduKateSG’s nine-category taxonomy is the optimal taxonomy.
That is what future research must test.
Conclusion
The Weak-Link Atlas begins with a simple change.
Do not ask only:
What did the child get wrong?
Ask:
What mechanism could have generated this wrong answer?
Then:
What evidence would distinguish that explanation from another one?
Then:
What intervention should change if our explanation is correct?
Then:
Did it?
That changes error correction into diagnostic science.
The wrong answer becomes:
observation
not:
identity
not:
final diagnosis
The proposed classes are:
Missing Node
Broken Edge
Weak Link
Wrong Edge
Routing
Translation
Transfer
Calibration
Regulation
But the most important principle sits above all nine:
A taxonomy is useful only if it changes intervention more accurately than a simpler alternative.
If a category cannot be distinguished:
merge it.
If it predicts nothing:
remove it.
If a new recurring failure cannot be represented:
investigate it.
If intervention does not produce the predicted downstream change:
revise the model.
That is the scientific function of PMRI-003.
Not to give eduKateSG more terminology.
To give eduKateSG a better way to be wrong, detect that it is wrong, and become more accurate.
Research Status
PMRI-001: Learning-System Framework
PMRI-002: Mathematics State Estimator
PMRI-003: Weak-Link Atlas v1.0
Next:
PMRI-004 — Mathematics Learning Continuity from Primary 1 to PSLE
Primary research question:
Which mathematical capabilities must remain sufficiently available and connected for later learning to remain efficient—and how can we distinguish a true prerequisite from a merely helpful earlier skill?
