The eduKateSG Mathematics State Estimator
PMRI-002 — Primary Mathematics Research Institute Series
eduKateSG Learning Sciences & Education Research Programme
Framework: Mathematics State Estimator MSE v0.1
Status: Measurement Framework / Research Protocol
Research Domain: Primary Mathematics / Formative Assessment / Diagnostic Education
Jurisdictional Context: Singapore Primary Mathematics, Primary 1–6
Last Reviewed: August 2026
Research Classification
Article Type: Measurement framework + evidence synthesis + proposed research protocol
Primary Research Question:
What observations are required before we can make a useful inference about what a child actually understands in Mathematics?
Secondary Research Question:
Can multiple low-cost observations collected during ordinary Mathematics teaching estimate learner state more accurately than score alone?
This paper does not claim that:
- eduKateSG currently possesses a validated psychometric diagnostic instrument;
- a tutor can read a child’s internal cognitive state directly;
- one response can reliably identify one underlying mechanism;
- the proposed observation dimensions should be combined into a universal Mathematics score.
This paper proposes:
a structured way of collecting and interpreting multiple observations before deciding what mathematical intervention should come next.
Start Here: https://edukatesg.com/how-mathematics-works/
Quick Read
A correct answer does not necessarily prove complete understanding.
A wrong answer does not necessarily prove absence of understanding.
A score tells us what happened on an assessment.
It does not automatically tell us why.
The eduKateSG Mathematics State Estimator therefore asks for multiple observations.
For example:
Accuracy
Did the learner get it right?
Latency
How readily did relevant Mathematics become available?
Explanation
Can the learner explain the mathematical relationship?
Representation
Can the same idea be shown another way?
Retrieval
Can the learner reproduce the Mathematics after time has passed?
Routing
Can the learner recognise when this Mathematics should be used?
Transfer
Can the learner still solve when the question changes?
Prompt Dependence
How much external help is required?
Calibration
Can the learner recognise and repair a suspicious answer?
Performance
Does the capability survive mixed and timed conditions?
Together, these observations produce a richer estimate.
The core rule is:
Do not diagnose from one signal when several different mechanisms can produce the same signal.
The proposed process is:
Observe → Triangulate → Form Hypothesis → Intervene → Retest → Delay → Transfer → Update State
This is the measurement layer underneath the eduKateSG Primary Mathematics research programme.
Why This Question Matters
Singapore’s current Assessment for Learning architecture is already moving explicitly in the direction of diagnosis rather than score-only interpretation.
SEAB describes MathsCheckPlus as helping teachers determine where students stand and plan targeted remediation before progression to the next level. CATalytics then allows deeper investigation of specific conceptual gaps. SEAB recommends using the tools together to obtain a fuller picture of mathematical readiness. (SEAB)
CATalytics reports do not stop at a test total. They include student and group profiles, profile descriptors and common-error analysis, and report learning outcomes that students can manage, partially manage or cannot manage. (SEAB)
That establishes an important principle:
A useful assessment should help change what happens next.
eduKateSG takes that principle into the small-group teaching environment and asks:
How much additional diagnostic resolution can we obtain from observing the learner while Mathematics is actually being performed?
Part 1 of 3
The Measurement Problem
What Does It Mean to “Know” Mathematics?
Consider a child who answers:
3/4 + 1/4 = 1
correctly.
What has been established?
At minimum:
The child produced the correct response to that task under those conditions.
But several stronger claims would require additional evidence.
Can the learner explain why the answer is 1?
Can the learner draw the relationship?
Can the learner solve:
5/8 + 3/8
without an example beside it?
Can the learner retrieve the method next week?
Can the learner recognise the same fraction structure inside a word problem?
Can the learner distinguish this from a question where denominators differ?
Can the learner do it without the tutor saying:
“Remember, the denominators are the same”?
These are different measurements.
So the first rule of the State Estimator is:
Correct Response ≠ Complete Understanding
Likewise:
Incorrect Response ≠ No Understanding
A child may understand the mathematical concept and make one arithmetic error.
That is very different from lacking the concept.
Observable Behaviour and Hidden State
We can directly observe:
- answers;
- working;
- time taken;
- explanations;
- diagrams;
- strategy choices;
- requests for help;
- corrections.
We cannot directly observe:
“understanding”
as an object.
Understanding is inferred from behaviour.
This distinction is fundamental.
The proposed State Estimator therefore contains two layers:
Observation Layer
What actually happened?
Inference Layer
What learner state best explains the pattern?
We must never confuse the second with the first.
The Observation Rule
A tutor should be able to write:
“The student solved three equivalent-fraction questions independently, required a prompt on a fourth representation, and could not retrieve the method five days later.”
That is observation.
A statement such as:
“The student has weak fraction intelligence.”
is neither sufficiently precise nor sufficiently grounded.
The research discipline is:
record behaviour before interpreting behaviour.
One Signal Can Have Several Explanations
Suppose the child takes a long time to solve:
8 × 7
Possible explanations include:
- multiplication retrieval is weak;
- attention was temporarily elsewhere;
- the learner was checking unusually carefully;
- anxiety slowed responding;
- the child normally reconstructs rather than retrieves the fact.
So:
High Latency
does not automatically equal:
Retrieval Deficit
Latency becomes informative only when interpreted alongside other evidence.
This is why we need triangulation.
Triangulation
Proposed Definition
Triangulation is the use of multiple observations that converge on, or contradict, a proposed learner-state explanation.
Suppose we hypothesise:
multiplication retrieval is unstable.
Useful supporting observations might include:
- slow multiplication facts;
- repeated reconstruction;
- correct answers when given sufficient time;
- improved speed after retrieval practice;
- downstream slowing in fraction problems.
If instead the child retrieves multiplication facts rapidly in isolation but fails only inside word problems, the original hypothesis weakens.
Perhaps the bottleneck is elsewhere.
That is desirable.
A diagnostic model should allow evidence to disconfirm it.
Formative Assessment Is a Process, Not One Event
IES guidance on formative assessment describes it as an ongoing, planned practice involving multiple sources of evidence gathered over time and used to adjust teaching. (ies.ed.gov)
This is a useful principle for eduKateSG.
The Mathematics State Estimator should not be:
one enormous diagnostic test at the start of the year.
It should operate continuously.
Every suitable learning event can update the estimate.
State Estimation Over Time
Conceptually:
Estimated State at t₁
↓
New Observation
↓
Updated State at t₂
↓
Intervention
↓
New Observation
↓
Updated State at t₃
The model is dynamic.
The child learns.
Forgets.
Reconnects.
Becomes faster.
Transfers.
Sometimes regresses temporarily.
Therefore:
diagnosis should be updateable.
A label that cannot change when the learner changes is a poor learning model.
Confidence Matters
A scientific state estimator should not pretend that every inference is equally certain.
eduKateSG therefore proposes three preliminary diagnostic confidence levels.
Low Confidence
One or two ambiguous observations.
Example:
One multiplication error appeared.
Do not overdiagnose.
Moderate Confidence
A repeated pattern appears across several tasks or representations.
Example:
The learner repeatedly reconstructs multiplication facts slowly across several sessions.
A working intervention hypothesis becomes reasonable.
Higher Confidence
Several independent observations converge and an intervention-response test supports the explanation.
Example:
slow multiplication retrieval
- ●
strong conceptual explanation
- ●
accurate untimed work
- ●
retrieval intervention improves speed
- ●
downstream fraction fluency improves
Now the retrieval hypothesis receives stronger support.
Even then, we should say:
evidence supports the diagnosis.
Not:
diagnosis is infallible.
The Minimum Evidence Principle
The purpose is not to collect maximum data.
It is to collect enough evidence to make the next instructional decision better.
So:
Diagnostic Resolution should be proportional to Decision Need.
A minor one-off error does not require a 40-minute investigation.
A persistent P5 fraction failure affecting several topics deserves more resolution.
This prevents research architecture from overwhelming teaching.
Research Question
What is the minimum observation set required to distinguish common Primary Mathematics failure mechanisms reliably enough to improve intervention selection?
This is a genuine future validation problem.
PMRI-002 proposes the candidate observations.
It does not yet claim the minimum set has been established.
Part 2 of 3
The Mathematics State Estimator
The proposed estimator observes ten major signals.
These signals correspond to, but are not identical with, the learning-state dimensions introduced in PMRI-001.
Sensor 1 — Accuracy
Question
Did the student produce mathematically correct output?
Accuracy remains essential.
Without it, we lose the most obvious performance signal.
But accuracy should be decomposed.
Was the:
method correct?
working correct?
calculation correct?
final answer correct?
unit correct?
quantity requested actually answered?
A binary:
right / wrong
often throws away useful information.
Accuracy Decomposition
Consider a 4-mark problem.
Case A
Correct model.
Correct method.
One arithmetic slip.
Wrong final answer.
Case B
Incorrect conceptual model from the beginning.
Several correct calculations follow from that wrong model.
Both may lose marks.
But their instructional implications differ.
The State Estimator therefore stores:
where correctness first diverged.
That first divergence is often more diagnostically valuable than the final answer.
Sensor 2 — Latency
Question
How much time or hesitation occurs before useful mathematical action begins?
Latency can reveal:
- retrieval friction;
- uncertainty;
- poor method selection;
- language difficulty;
- unfamiliarity.
But latency must be contextualised.
Fast is not automatically good.
A student can answer quickly and incorrectly because of impulsive routing.
Slow is not automatically weak.
A student may be evaluating several strategies.
So latency is not a grade.
It is a sensor.
Latency Map
A useful distinction is:
Retrieval Latency
How long before a known fact becomes available?
Routing Latency
How long before the learner identifies what type of Mathematics is relevant?
Execution Latency
How long does the actual procedure take after the route is known?
These may reveal different bottlenecks.
Sensor 3 — Explanation
Question
Can the learner communicate the mathematical relationship underneath the procedure?
Explanation can distinguish:
procedure reproduction
from:
structural understanding
For example:
Why does multiplying length by breadth give rectangle area?
or:
Why are 2/4 and 1/2 equal?
The student does not need university-level mathematical language.
We are looking for evidence that the learner sees the relevant relationship.
Explanation Should Not Become a Language Trap
There is a caution.
A learner may understand mathematically but have difficulty expressing the idea verbally.
Therefore:
poor verbal explanation ≠ automatically poor mathematical understanding.
We can ask for another form.
Draw it.
Show it with numbers.
Build an example.
Demonstrate it.
This connects explanation with representation.
Sensor 4 — Representation Flexibility
Question
Can the learner move between different representations of the same mathematical relationship?
Examples:
objects ↔ drawing ↔ number sentence
fraction ↔ diagram
fraction ↔ decimal ↔ percentage
word problem ↔ bar model
table ↔ graph
3D structure ↔ layers of unit cubes
Representation is central to mathematical problem solving, and evidence-based IES guidance supports deliberate use of visual representations and multiple strategies. (ies.ed.gov)
The State Estimator therefore treats representation as something to observe rather than merely something the tutor supplies.
Representation Probe
After a correct symbolic answer, ask:
Can you show this another way?
If yes, confidence in structural understanding may increase.
If the learner can only reproduce the exact trained format, transfer may be more fragile.
Sensor 5 — Prompt Dependence
Question
How much assistance is required before successful mathematical action occurs?
This is one of the most important tuition-specific observations.
Consider four states.
P0 — Full Support
Tutor demonstrates most of the solution.
P1 — Strategic Prompt
Tutor tells the learner what method or representation to use.
P2 — Directional Prompt
Tutor asks a question that redirects attention.
P3 — Independent
Student initiates and completes the route independently.
A child moving from P0 toward P3 is changing state even before examination marks show a dramatic increase.
Prompt Dependence Can Hide Behind Correct Answers
Suppose two students both complete a worksheet correctly.
Student A needed:
“Draw a model.”
“Now find the total.”
“Remember percentage.”
Student B selected all of those actions independently.
The completed worksheet may look similar.
The learner states are not.
This is one reason a very small teaching environment can reveal information that a final answer does not.
The Prompt-Fading Test
The intervention sequence should deliberately test:
help
↓
less help
↓
no help
If correctness collapses every time support disappears, repair is incomplete.
The goal is not zero teaching.
The goal is to measure how much control has transferred to the learner.
Sensor 6 — Delayed Retrieval
Question
Does the learning remain accessible after time passes?
Immediate success may partly reflect:
- recent explanation;
- current worksheet pattern;
- working memory of the previous example.
Delayed retrieval weakens those supports.
Ask again:
tomorrow
or:
next week
or:
after other topics have intervened
If the Mathematics returns independently, evidence of durable availability becomes stronger.
Research on mathematics learning distinguishes fluency and transfer as related but different outcomes, and IES-funded work has specifically investigated how arithmetic practice produces different patterns of fluent and transferable knowledge. (ies.ed.gov)
Retrieval State Levels
A preliminary eduKateSG scale might be:
R0 — Not Retrievable
Cannot reconstruct even with substantial support.
R1 — Recognition
Method becomes familiar when shown.
R2 — Cued Retrieval
Can retrieve with a hint.
R3 — Independent Retrieval
Can retrieve without hint.
R4 — Integrated Retrieval
Can retrieve while solving a different or mixed problem.
These are proposed operational categories.
They are not yet validated psychometric levels.
Sensor 7 — Routing
Question
Can the learner identify which Mathematics is relevant when the method is not announced?
This sensor becomes increasingly important from P3 onward.
Blocked worksheet:
Percentage Practice
has already provided routing information.
Mixed assessment:
no topic label.
Now the child must discriminate.
We can test routing by placing several mathematical structures close together.
For example:
fraction
area
percentage
division
ratio
Then observe whether the learner identifies the relevant family before calculating.
Routing Error vs Knowledge Error
Suppose a learner selects division when multiplication is required.
Before reteaching multiplication, ask:
Does the child know multiplication?
If yes, the problem may be method selection.
This distinction is critical.
A student can possess the necessary Mathematics and still route incorrectly.
That requires a different intervention.
Sensor 8 — Transfer
Question
Does the learner retain control when the surface of the problem changes?
Transfer should be tested systematically.
A useful ladder is:
T0 — Exact Reproduction
Same structure and very similar surface.
T1 — Numerical Variation
Numbers change.
T2 — Language Variation
Wording changes.
T3 — Representation Variation
Diagram or mathematical form changes.
T4 — Context Variation
Story changes.
T5 — Integrated Transfer
Concept appears alongside another concept.
T6 — Unfamiliar Problem
No obvious stored template is available.
Again, these are proposed research levels rather than validated scales.
Their purpose is to give the research programme something explicit to test.
Why Transfer Matters
IES notes that Mathematics students often struggle to apply learning in new contexts and has funded research specifically examining which instructional conditions produce more fluent and transferable arithmetic learning. (ies.ed.gov)
So the State Estimator should not ask only:
Can the student repeat?
It should ask:
How far can the learning travel before it collapses?
That is a powerful measurement question.
Sensor 9 — Error Recurrence
Question
Does the same failure mechanism return?
One wrong answer may be noise.
Repeated related errors become a signal.
Examples:
- repeated unit loss;
- recurring denominator mistakes;
- repeated wrong-operation selection;
- repeated percentage reference errors;
- repeated unfinished multi-step questions.
The useful measurement is not merely:
number of errors.
It is:
error genealogy.
Which errors are descendants of the same underlying mechanism?
This becomes central to PMRI-003.
Repair Should Change Error Recurrence
If an intervention was successful, we expect some future behaviour to change.
For example:
Before repair
same fraction-equivalence error repeatedly appears.
After repair
the error should:
- disappear;
- decrease substantially;
- or change form in a way suggesting partial progress.
If the exact error continues unchanged, the intervention should be questioned.
This is the receiver test again.
Sensor 10 — Calibration and Self-Correction
Question
Can the learner detect when something is wrong without the tutor announcing it?
Examples:
A pencil is 350 metres long.
25% of a number is larger than 500% of the same positive number.
The area is written in centimetres rather than square centimetres.
A learner with stronger calibration may pause and investigate.
This is important because examinations cannot provide immediate corrective feedback.
The learner needs internal monitoring.
Self-Correction Is Strong Evidence
Consider two wrong responses.
Learner A
Produces wrong answer.
Accepts it immediately.
Learner B
Produces wrong answer.
Says:
“That doesn’t make sense.”
Returns to working.
Repairs it.
Both temporarily made an error.
Only one demonstrated effective error regulation.
So the State Estimator should record:
error recovery
not only error occurrence.
Sensor 11 — Performance Under Constraint
The preceding sensors examine capability.
Eventually we need to ask:
Does the capability survive load?
Possible constraints include:
- mixed topics;
- limited time;
- no tutor prompts;
- longer papers;
- unfamiliar ordering;
- examination pressure.
A learner can possess a capability in isolation but fail to dispatch it under load.
This is particularly important in P5, P6 and PSLE Mathematics.
Capability State vs Performance State
We therefore distinguish:
Capability State
What the learner can do under sufficiently supportive conditions.
Dispatchability State
How readily that capability can be brought online.
Performance State
How reliably it survives realistic constraints.
This prevents a common error:
assuming examination failure always means the Mathematics was never understood.
Sometimes it was understood but not reliably operational.
The State Estimator Matrix
The proposed observational matrix is:
| Sensor | Core question |
|---|---|
| Accuracy | Was the Mathematics correct? |
| Latency | How readily did useful action begin? |
| Explanation | Can the relationship be communicated? |
| Representation | Can the idea be transformed? |
| Prompt dependence | How much help is required? |
| Delayed retrieval | Does the learning survive time? |
| Routing | Can the relevant Mathematics be selected? |
| Transfer | Does learning survive surface change? |
| Error recurrence | Does the same mechanism return? |
| Calibration | Can error be detected and repaired? |
| Constraint performance | Does capability survive mixed/timed conditions? |
The key point is:
No single column is Mathematics understanding.
Understanding is inferred from the pattern.
A Worked State-Estimator Example
Suppose a Primary 5 learner is reportedly:
“weak at percentage.”
We run several observations.
Observation 1
Direct percentage calculation:
80% of 250
Correct.
Observation 2
Student explains that percent means:
“out of one hundred.”
Reasonable conceptual evidence.
Observation 3
Can convert:
1/4 → 25%
Correct.
Observation 4
Word problem involving percentage increase.
Wrong.
Observation 5
Tutor says:
“What is the original 100%?”
Student immediately reconstructs the problem correctly.
Observation 6
New percentage increase problem without hint.
Wrong again.
What does this suggest?
Not necessarily:
percentage knowledge missing.
Instead, evidence may support:
routing/reference-quantity control is unstable and prompt dependence remains high.
So the intervention changes.
Less:
repeat direct percentage calculations.
More:
discriminate reference quantities across mixed percentage situations.
That is what state estimation is for.
Another Example: Same Score, Different Intervention
Two students score 6/10 on a fraction probe.
Student A
- understands fraction diagrams;
- explains equivalence;
- retrieves slowly;
- makes multiplication errors.
Likely intervention:
improve low-level retrieval/execution.
Student B
- calculates familiar questions correctly;
- cannot explain equivalence;
- fails representation changes;
- depends on memorised procedure.
Likely intervention:
rebuild conceptual/representation structure.
Same score:
6/10
Different estimated state.
Different repair.
Part 3 of 3
From Observation to Research Method
The State Estimator must not become tutor intuition dressed in technical language.
It needs a disciplined protocol.
Step 1 — Define the Target Capability
Do not assess:
Mathematics.
Assess something narrower.
For example:
equivalent fractions.
or:
multiplicative comparison.
or:
selecting between area and perimeter.
The more precise the target, the more interpretable the observations.
Step 2 — Establish a Baseline
Collect enough information to answer:
What can the learner currently do without intervention?
Use appropriately chosen tasks.
Record:
- response;
- method;
- latency where meaningful;
- representation;
- prompts;
- confidence;
- self-correction.
This becomes the baseline state estimate.
Step 3 — Vary One Important Dimension
A good diagnostic probe changes something.
For example:
same concept, different number
or:
same concept, different representation
or:
same concept, different context
The change helps reveal what part of the previous success was robust.
Step 4 — Remove Support
If the student succeeds after tutoring:
reduce prompting.
This tests ownership.
Success under full scaffolding should not be reported as independent mastery.
Step 5 — Delay
Do not conclude repair from the immediate post-test alone.
Return after time.
This measures availability after recent instructional support has faded.
Step 6 — Mix
Place the concept among alternatives.
This tests routing.
The student now has to decide which mathematical system is relevant.
Step 7 — Transfer
Change the surface.
If the learning survives, confidence in the state estimate increases.
Step 8 — Apply Constraint
For older Primary students, test whether the capability survives:
- time;
- mixed papers;
- longer working sequences.
This moves from learning measurement toward performance measurement.
Step 9 — Update the Hypothesis
The estimated state should change when evidence changes.
Suppose we initially believed:
concept weak.
After several probes we discover:
- explanation strong;
- representation strong;
- untimed accuracy high;
- timed performance poor.
The hypothesis should change toward:
execution/regulation bottleneck.
A research institution should reward corrected diagnosis.
Not defend its first guess.
The Intervention Test
One of the strongest ways to evaluate a diagnostic hypothesis is to intervene.
Suppose we believe:
multiplication retrieval is constraining fraction performance.
Prediction:
improving multiplication retrieval should reduce fraction-task latency or error under otherwise similar conditions.
If multiplication improves but fraction performance remains unchanged, the weak-link hypothesis receives less support.
Maybe the bottleneck lies elsewhere.
This turns intervention into a diagnostic experiment.
Diagnosis Should Make Predictions
A useful state estimate should predict something.
For example:
Hypothesis
Representation is the bottleneck.
Prediction
When supplied an appropriate diagram, performance should improve materially.
If not:
reconsider the hypothesis.
Or:
Hypothesis
Knowledge exists but retrieval is weak.
Prediction
Recognition will be substantially stronger than unaided delayed recall.
If recognition and recall are both absent:
the concept itself may be less established than expected.
Predictions make the system testable.
Competing Hypotheses
When evidence is ambiguous, keep multiple explanations alive.
Example:
Student fails two-step problems.
Candidate explanations:
H1 — sequencing weakness
H2 — multiplication retrieval bottleneck
H3 — translation weakness
H4 — regulation under load
Now design probes that distinguish them.
That is more scientific than immediately selecting whichever explanation sounds most plausible.
Falsification Condition
Each diagnosis should have a condition under which it weakens.
Example:
Proposed diagnosis
Routing weakness.
Weakening evidence
Student selects appropriate methods reliably across several mixed tasks without prompting.
Then routing is probably not the main bottleneck.
Move elsewhere.
This prevents categories from becoming unfalsifiable explanations for everything.
No Universal State Score Yet
It may be tempting to take the eleven sensors and produce:
Mathematics State = 82.7
We should not do this yet.
There are several unanswered questions:
- Are the sensors independent?
- Which are task-specific?
- Which are age-specific?
- How reliable are tutor judgements?
- How should latency be normalised?
- Is transfer one dimension or several?
- How should prompt dependence be weighted?
- Does one combined score improve intervention?
Until those questions are studied:
preserve the profile.
Do not compress prematurely.
The Mathematics State Profile
A future parent-facing report could look conceptually like:
Fraction Concept
Stable
Retrieval
Moderately stable
Representation
Stable
Routing
Needs support
Transfer
Unstable
Prompt Dependence
Moderate
Error Recurrence
High in mixed problems
Suggested Next Intervention
Mixed discrimination + delayed transfer retest
This is more useful than inventing:
Fraction Intelligence Score: 67.
Observation Reliability
A research institution also needs to ask:
Would two trained tutors interpret the same performance similarly?
This is an important future validation requirement.
The State Estimator cannot depend entirely on one tutor’s subjective impression.
Potential research steps include:
- explicit observation definitions;
- example responses;
- coding rubrics;
- double coding;
- agreement checks;
- adjudication of ambiguous cases.
This will become necessary before stronger scientific claims can be made.
Measurement Validity
Another question:
Are we measuring what we think we are measuring?
Suppose latency is used as a retrieval measure.
But a child is naturally deliberate and slow across all tasks.
Then latency may partly capture response style rather than mathematical retrieval.
So each sensor needs validity testing.
This is exactly why the present framework is labelled:
MSE v0.1
It is a proposed measurement system.
Not a finished one.
State Estimator Research Agenda
The first validation programme should investigate:
RQ1
How reliably can tutors code the proposed observations?
RQ2
Which sensors provide information beyond score alone?
RQ3
Which sensor combinations best predict independent performance?
RQ4
Which observations predict delayed retention?
RQ5
Which observations predict transfer?
RQ6
Can prompt dependence predict later independent performance?
RQ7
Can error recurrence distinguish weak-link types?
RQ8
Does routing performance on mixed tasks predict examination performance better than blocked-practice accuracy?
RQ9
How stable are state estimates across days?
RQ10
Which state dimensions change most after targeted intervention?
These are empirical questions.
They form part of the institution’s future research programme.
Singapore Alignment Without Overclaiming
SEAB’s adaptive Mathematics tools provide an important external precedent for going beyond a single total score.
MathsCheckPlus is designed to identify readiness and support targeted remediation, while CATalytics provides more focused diagnostic information on prior knowledge and learning gaps. (SEAB)
CATalytics currently assesses P5/P6 proficiency in selected Standard Mathematics topics and supplies student profiles, descriptors and common-error information. (SEAB)
eduKateSG should be precise about the relationship.
We are not claiming:
eduKateSG’s State Estimator is the same as SEAB’s assessment system.
Nor:
SEAB validates our specific categories.
What SEAB does validate is the importance of diagnostic resolution:
understanding where the learning gap lies can support more targeted intervention than score alone.
Our research question begins from there.
What a Small-Group Environment Adds
Standardised assessments have enormous strengths:
- consistency;
- coverage;
- comparability;
- scalable reporting.
A three-student tuition environment offers different potential information.
The tutor can observe:
- strategy before the answer;
- hesitation;
- self-talk;
- representation choice;
- response to prompts;
- self-correction;
- what happens when support is removed.
So the two environments need not compete.
They answer different measurement questions.
The research opportunity is:
Can high-resolution observational evidence improve the instructional usefulness of formal assessment evidence?
That is a much more defensible institutional research question.
The Lesson as a Repeated Measurement Environment
A normal lesson contains many natural probes.
Student attempts independently.
Observation.
Tutor asks a strategic question.
Observation.
Student changes representation.
Observation.
Hint is removed.
Observation.
Topic returns next week.
Observation.
Same Mathematics appears in a mixed paper.
Observation.
This means measurement can occur inside teaching without requiring constant formal testing.
But only if the observations are recorded systematically enough to become meaningful.
Do Not Turn the Child Into a Dashboard
This is an important boundary.
More measurement is not always better.
A learner is not:
a collection of performance indicators.
The purpose of the State Estimator is to make intervention more precise while keeping teaching human, calm and proportionate.
So:
Measurement serves learning.
Learning does not serve measurement.
Parent Communication
A research-oriented tuition organisation should also improve how findings are communicated to parents.
Instead of:
“Your child is weak in Maths.”
report:
“Whole-number calculation is generally stable. Fraction concepts are understood when represented visually, but delayed retrieval and mixed-question selection remain inconsistent. We are therefore working on retrieval and routing before increasing paper volume.”
This is:
- more precise;
- more actionable;
- less identity-forming.
It also tells the parent what is being tested next.
Student Communication
For the learner:
Instead of:
“You don’t know fractions.”
try:
“You understand the fraction when we draw it. The part we are strengthening now is recognising when to use that idea without a hint.”
This changes the psychological meaning of diagnosis.
The weakness becomes:
a specific trainable state.
Not:
a verdict about mathematical ability.
Tutor Communication
The tutor’s internal record should answer:
What was observed?
↓
What hypothesis does it support?
↓
How confident are we?
↓
What alternative explanation remains?
↓
What intervention follows?
↓
What result would confirm or weaken the hypothesis?
This is the beginning of research discipline inside ordinary practice.
The MSE v0.1 Observation Record
A minimal future record could contain:
Target
What Mathematics is being observed?
Task
What did the learner attempt?
Accuracy
What happened?
First Divergence
Where did the reasoning first leave the correct path?
Latency
Was response unusually immediate, normal or delayed?
Representation
What representation was selected?
Prompt Level
How much support was required?
Explanation
What relationship could the learner articulate?
Retrieval
Was the Mathematics independently available?
Transfer
Did it survive variation?
Self-Correction
Did the learner detect error?
Current Hypothesis
What mechanism currently best explains the pattern?
Confidence
Low / Moderate / Higher
Competing Hypothesis
What alternative still fits?
Intervention
What will we change?
Retest
What result will tell us whether the intervention worked?
This is enough structure to make the process inspectable without turning every lesson into laboratory bureaucracy.
Evidence Boundary
What Current Evidence Supports
Current assessment and education research strongly supports using assessment evidence to inform instruction rather than treating assessment only as an endpoint. SEAB’s adaptive Mathematics tools explicitly identify learning gaps and provide targeted information for remediation. (SEAB)
Research also treats mathematical fluency and transfer as distinguishable outcomes and continues to investigate instructional conditions that help learners apply arithmetic knowledge beyond trained problems. (ies.ed.gov)
What PMRI-002 Adds
eduKateSG proposes a higher-resolution observational framework using:
accuracy
latency
explanation
representation
prompt dependence
delayed retrieval
routing
transfer
error recurrence
calibration
performance under constraint
These dimensions are research constructs requiring validation.
What We Do Not Yet Know
We do not yet know:
- which observations are redundant;
- which observations are most predictive;
- how reliable tutor coding will be;
- whether eleven sensors are too many;
- whether some sensors should split;
- whether some should merge;
- which measures work best at different Primary levels;
- whether the resulting state estimates improve learning outcomes compared with simpler diagnostic approaches.
Those questions are not weaknesses in the research programme.
They are the research programme.
The eduKateSG Mathematics State Estimator
The complete MSE v0.1 runtime is:
Define Target
What mathematical capability are we investigating?
↓
Observe
What does the learner actually do?
↓
Decompose
Where does performance first diverge?
↓
Sample Multiple Sensors
Accuracy
Latency
Explanation
Representation
Prompt Dependence
Retrieval
Routing
Transfer
Error Recurrence
Calibration
Constraint Performance
↓
Triangulate
Which observations converge?
↓
Generate Hypotheses
What mechanisms could explain the pattern?
↓
Assign Confidence
Low / Moderate / Higher
↓
Keep Alternatives
What else could explain it?
↓
Intervene
Change one useful variable.
↓
Immediate Retest
Did behaviour change?
↓
Remove Support
Does success remain?
↓
Delay
Does the learning return later?
↓
Mix
Can the learner route correctly?
↓
Transfer
Does learning survive surface change?
↓
Constraint Test
Does capability remain available under load?
↓
Update State
What does the new evidence support?
↓
Record
What did we learn?
↓
Research Memory
Does the case support or challenge the framework?
From “How Many Marks?” to “What State Produced the Marks?”
Marks remain important.
The State Estimator does not replace them.
It gives us a way to ask the next question.
A student scores:
68%
We ask:
What mathematical state generated the 68%?
Another student scores:
68%
We ask again.
The answer may be different.
That difference is where targeted education begins.
Conclusion
Understanding Must Be Inferred Carefully
There is no single worksheet question that reveals everything a child understands.
There is no single mark that fully describes the learner.
There is no single wrong answer that identifies one unique cause.
So the research problem is:
How do we infer the learner’s state without pretending that our inference is the learner?
The eduKateSG Mathematics State Estimator begins with a simple discipline:
observe first
↓
infer cautiously
↓
triangulate
↓
test the inference
↓
change the intervention
↓
observe again
That is measurement as a learning loop.
Not measurement as labelling.
The child may:
recognise but not retrieve
retrieve but not route
route but execute poorly
execute correctly but fail transfer
transfer without time pressure but collapse during examination conditions
Each state produces a different educational problem.
And therefore potentially a different repair.
That is why:
Score ≠ State
and:
State ≠ Identity
The goal is not to construct the most complicated diagnostic system possible.
It is to obtain just enough resolution to make the next educational decision better.
When the evidence is weak:
remain uncertain.
When several observations converge:
form a stronger hypothesis.
When the intervention fails:
revise the diagnosis.
When transfer fails:
do not declare mastery.
When delayed retrieval fails:
do not confuse immediate performance with durable learning.
When support is still required:
do not report independent control.
And when the learner changes:
update the state.
That is the Mathematics State Estimator.
Not a final test.
A continuously improving estimate of:
What can this learner currently do, under what conditions, with how much support, and what needs to change next?
That is the measurement architecture needed before eduKateSG can credibly build the next research layer.
Research Status at Publication
PMRI-001
Primary Mathematics Learning-System Framework established.
PMRI-002
Mathematics State Estimator v0.1 established as a proposed measurement framework.
Next Research Paper:
PMRI-003 — The Primary Mathematics Weak-Link Atlas
Why the Same Wrong Answer Can Have Different Causes
Primary research question:
Can recurring mathematical failure patterns be classified into a small enough set of diagnostic categories to improve intervention selection without oversimplifying the learner?
Candidate failure classes:
Missing Node
Broken Edge
Weak Link
Wrong Edge
Routing
Translation
Transfer
Calibration
Regulation
PMRI-003 will attempt to convert those categories from useful language into a falsifiable diagnostic taxonomy.
