VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Do We Know What a Child Actually Understands in Mathematics?

The eduKateSG Mathematics State Estimator

PMRI-002 — Primary Mathematics Research Institute Series
eduKateSG Learning Sciences & Education Research Programme
Framework: Mathematics State Estimator MSE v0.1
Status: Measurement Framework / Research Protocol
Research Domain: Primary Mathematics / Formative Assessment / Diagnostic Education
Jurisdictional Context: Singapore Primary Mathematics, Primary 1–6
Last Reviewed: August 2026


Research Classification

Article Type: Measurement framework + evidence synthesis + proposed research protocol

Primary Research Question:

What observations are required before we can make a useful inference about what a child actually understands in Mathematics?

Secondary Research Question:

Can multiple low-cost observations collected during ordinary Mathematics teaching estimate learner state more accurately than score alone?

This paper does not claim that:

  • eduKateSG currently possesses a validated psychometric diagnostic instrument;
  • a tutor can read a child’s internal cognitive state directly;
  • one response can reliably identify one underlying mechanism;
  • the proposed observation dimensions should be combined into a universal Mathematics score.

This paper proposes:

a structured way of collecting and interpreting multiple observations before deciding what mathematical intervention should come next.

Start Here: https://edukatesg.com/how-mathematics-works/


Quick Read

A correct answer does not necessarily prove complete understanding.

A wrong answer does not necessarily prove absence of understanding.

A score tells us what happened on an assessment.

It does not automatically tell us why.

The eduKateSG Mathematics State Estimator therefore asks for multiple observations.

For example:

Accuracy

Did the learner get it right?

Latency

How readily did relevant Mathematics become available?

Explanation

Can the learner explain the mathematical relationship?

Representation

Can the same idea be shown another way?

Retrieval

Can the learner reproduce the Mathematics after time has passed?

Routing

Can the learner recognise when this Mathematics should be used?

Transfer

Can the learner still solve when the question changes?

Prompt Dependence

How much external help is required?

Calibration

Can the learner recognise and repair a suspicious answer?

Performance

Does the capability survive mixed and timed conditions?

Together, these observations produce a richer estimate.

The core rule is:

Do not diagnose from one signal when several different mechanisms can produce the same signal.

The proposed process is:

Observe → Triangulate → Form Hypothesis → Intervene → Retest → Delay → Transfer → Update State

This is the measurement layer underneath the eduKateSG Primary Mathematics research programme.


Why This Question Matters

Singapore’s current Assessment for Learning architecture is already moving explicitly in the direction of diagnosis rather than score-only interpretation.

SEAB describes MathsCheckPlus as helping teachers determine where students stand and plan targeted remediation before progression to the next level. CATalytics then allows deeper investigation of specific conceptual gaps. SEAB recommends using the tools together to obtain a fuller picture of mathematical readiness. (SEAB)

CATalytics reports do not stop at a test total. They include student and group profiles, profile descriptors and common-error analysis, and report learning outcomes that students can manage, partially manage or cannot manage. (SEAB)

That establishes an important principle:

A useful assessment should help change what happens next.

eduKateSG takes that principle into the small-group teaching environment and asks:

How much additional diagnostic resolution can we obtain from observing the learner while Mathematics is actually being performed?


Part 1 of 3

The Measurement Problem

What Does It Mean to “Know” Mathematics?

Consider a child who answers:

3/4 + 1/4 = 1

correctly.

What has been established?

At minimum:

The child produced the correct response to that task under those conditions.

But several stronger claims would require additional evidence.

Can the learner explain why the answer is 1?

Can the learner draw the relationship?

Can the learner solve:

5/8 + 3/8

without an example beside it?

Can the learner retrieve the method next week?

Can the learner recognise the same fraction structure inside a word problem?

Can the learner distinguish this from a question where denominators differ?

Can the learner do it without the tutor saying:

“Remember, the denominators are the same”?

These are different measurements.

So the first rule of the State Estimator is:

Correct Response ≠ Complete Understanding

Likewise:

Incorrect Response ≠ No Understanding

A child may understand the mathematical concept and make one arithmetic error.

That is very different from lacking the concept.


Observable Behaviour and Hidden State

We can directly observe:

  • answers;
  • working;
  • time taken;
  • explanations;
  • diagrams;
  • strategy choices;
  • requests for help;
  • corrections.

We cannot directly observe:

“understanding”

as an object.

Understanding is inferred from behaviour.

This distinction is fundamental.

The proposed State Estimator therefore contains two layers:

Observation Layer

What actually happened?

Inference Layer

What learner state best explains the pattern?

We must never confuse the second with the first.


The Observation Rule

A tutor should be able to write:

“The student solved three equivalent-fraction questions independently, required a prompt on a fourth representation, and could not retrieve the method five days later.”

That is observation.

A statement such as:

“The student has weak fraction intelligence.”

is neither sufficiently precise nor sufficiently grounded.

The research discipline is:

record behaviour before interpreting behaviour.


One Signal Can Have Several Explanations

Suppose the child takes a long time to solve:

8 × 7

Possible explanations include:

  • multiplication retrieval is weak;
  • attention was temporarily elsewhere;
  • the learner was checking unusually carefully;
  • anxiety slowed responding;
  • the child normally reconstructs rather than retrieves the fact.

So:

High Latency

does not automatically equal:

Retrieval Deficit

Latency becomes informative only when interpreted alongside other evidence.

This is why we need triangulation.


Triangulation

Proposed Definition

Triangulation is the use of multiple observations that converge on, or contradict, a proposed learner-state explanation.

Suppose we hypothesise:

multiplication retrieval is unstable.

Useful supporting observations might include:

  • slow multiplication facts;
  • repeated reconstruction;
  • correct answers when given sufficient time;
  • improved speed after retrieval practice;
  • downstream slowing in fraction problems.

If instead the child retrieves multiplication facts rapidly in isolation but fails only inside word problems, the original hypothesis weakens.

Perhaps the bottleneck is elsewhere.

That is desirable.

A diagnostic model should allow evidence to disconfirm it.


Formative Assessment Is a Process, Not One Event

IES guidance on formative assessment describes it as an ongoing, planned practice involving multiple sources of evidence gathered over time and used to adjust teaching. (ies.ed.gov)

This is a useful principle for eduKateSG.

The Mathematics State Estimator should not be:

one enormous diagnostic test at the start of the year.

It should operate continuously.

Every suitable learning event can update the estimate.


State Estimation Over Time

Conceptually:

Estimated State at t₁

New Observation

Updated State at t₂

Intervention

New Observation

Updated State at t₃

The model is dynamic.

The child learns.

Forgets.

Reconnects.

Becomes faster.

Transfers.

Sometimes regresses temporarily.

Therefore:

diagnosis should be updateable.

A label that cannot change when the learner changes is a poor learning model.


Confidence Matters

A scientific state estimator should not pretend that every inference is equally certain.

eduKateSG therefore proposes three preliminary diagnostic confidence levels.

Low Confidence

One or two ambiguous observations.

Example:

One multiplication error appeared.

Do not overdiagnose.


Moderate Confidence

A repeated pattern appears across several tasks or representations.

Example:

The learner repeatedly reconstructs multiplication facts slowly across several sessions.

A working intervention hypothesis becomes reasonable.


Higher Confidence

Several independent observations converge and an intervention-response test supports the explanation.

Example:

slow multiplication retrieval

strong conceptual explanation

accurate untimed work

retrieval intervention improves speed

downstream fraction fluency improves

Now the retrieval hypothesis receives stronger support.

Even then, we should say:

evidence supports the diagnosis.

Not:

diagnosis is infallible.


The Minimum Evidence Principle

The purpose is not to collect maximum data.

It is to collect enough evidence to make the next instructional decision better.

So:

Diagnostic Resolution should be proportional to Decision Need.

A minor one-off error does not require a 40-minute investigation.

A persistent P5 fraction failure affecting several topics deserves more resolution.

This prevents research architecture from overwhelming teaching.


Research Question

What is the minimum observation set required to distinguish common Primary Mathematics failure mechanisms reliably enough to improve intervention selection?

This is a genuine future validation problem.

PMRI-002 proposes the candidate observations.

It does not yet claim the minimum set has been established.


Part 2 of 3

The Mathematics State Estimator

The proposed estimator observes ten major signals.

These signals correspond to, but are not identical with, the learning-state dimensions introduced in PMRI-001.


Sensor 1 — Accuracy

Question

Did the student produce mathematically correct output?

Accuracy remains essential.

Without it, we lose the most obvious performance signal.

But accuracy should be decomposed.

Was the:

method correct?

working correct?

calculation correct?

final answer correct?

unit correct?

quantity requested actually answered?

A binary:

right / wrong

often throws away useful information.


Accuracy Decomposition

Consider a 4-mark problem.

Case A

Correct model.

Correct method.

One arithmetic slip.

Wrong final answer.

Case B

Incorrect conceptual model from the beginning.

Several correct calculations follow from that wrong model.

Both may lose marks.

But their instructional implications differ.

The State Estimator therefore stores:

where correctness first diverged.

That first divergence is often more diagnostically valuable than the final answer.


Sensor 2 — Latency

Question

How much time or hesitation occurs before useful mathematical action begins?

Latency can reveal:

  • retrieval friction;
  • uncertainty;
  • poor method selection;
  • language difficulty;
  • unfamiliarity.

But latency must be contextualised.

Fast is not automatically good.

A student can answer quickly and incorrectly because of impulsive routing.

Slow is not automatically weak.

A student may be evaluating several strategies.

So latency is not a grade.

It is a sensor.


Latency Map

A useful distinction is:

Retrieval Latency

How long before a known fact becomes available?

Routing Latency

How long before the learner identifies what type of Mathematics is relevant?

Execution Latency

How long does the actual procedure take after the route is known?

These may reveal different bottlenecks.


Sensor 3 — Explanation

Question

Can the learner communicate the mathematical relationship underneath the procedure?

Explanation can distinguish:

procedure reproduction

from:

structural understanding

For example:

Why does multiplying length by breadth give rectangle area?

or:

Why are 2/4 and 1/2 equal?

The student does not need university-level mathematical language.

We are looking for evidence that the learner sees the relevant relationship.


Explanation Should Not Become a Language Trap

There is a caution.

A learner may understand mathematically but have difficulty expressing the idea verbally.

Therefore:

poor verbal explanation ≠ automatically poor mathematical understanding.

We can ask for another form.

Draw it.

Show it with numbers.

Build an example.

Demonstrate it.

This connects explanation with representation.


Sensor 4 — Representation Flexibility

Question

Can the learner move between different representations of the same mathematical relationship?

Examples:

objects ↔ drawing ↔ number sentence

fraction ↔ diagram

fraction ↔ decimal ↔ percentage

word problem ↔ bar model

table ↔ graph

3D structure ↔ layers of unit cubes

Representation is central to mathematical problem solving, and evidence-based IES guidance supports deliberate use of visual representations and multiple strategies. (ies.ed.gov)

The State Estimator therefore treats representation as something to observe rather than merely something the tutor supplies.


Representation Probe

After a correct symbolic answer, ask:

Can you show this another way?

If yes, confidence in structural understanding may increase.

If the learner can only reproduce the exact trained format, transfer may be more fragile.


Sensor 5 — Prompt Dependence

Question

How much assistance is required before successful mathematical action occurs?

This is one of the most important tuition-specific observations.

Consider four states.

P0 — Full Support

Tutor demonstrates most of the solution.

P1 — Strategic Prompt

Tutor tells the learner what method or representation to use.

P2 — Directional Prompt

Tutor asks a question that redirects attention.

P3 — Independent

Student initiates and completes the route independently.

A child moving from P0 toward P3 is changing state even before examination marks show a dramatic increase.


Prompt Dependence Can Hide Behind Correct Answers

Suppose two students both complete a worksheet correctly.

Student A needed:

“Draw a model.”

“Now find the total.”

“Remember percentage.”

Student B selected all of those actions independently.

The completed worksheet may look similar.

The learner states are not.

This is one reason a very small teaching environment can reveal information that a final answer does not.


The Prompt-Fading Test

The intervention sequence should deliberately test:

help

less help

no help

If correctness collapses every time support disappears, repair is incomplete.

The goal is not zero teaching.

The goal is to measure how much control has transferred to the learner.


Sensor 6 — Delayed Retrieval

Question

Does the learning remain accessible after time passes?

Immediate success may partly reflect:

  • recent explanation;
  • current worksheet pattern;
  • working memory of the previous example.

Delayed retrieval weakens those supports.

Ask again:

tomorrow

or:

next week

or:

after other topics have intervened

If the Mathematics returns independently, evidence of durable availability becomes stronger.

Research on mathematics learning distinguishes fluency and transfer as related but different outcomes, and IES-funded work has specifically investigated how arithmetic practice produces different patterns of fluent and transferable knowledge. (ies.ed.gov)


Retrieval State Levels

A preliminary eduKateSG scale might be:

R0 — Not Retrievable

Cannot reconstruct even with substantial support.

R1 — Recognition

Method becomes familiar when shown.

R2 — Cued Retrieval

Can retrieve with a hint.

R3 — Independent Retrieval

Can retrieve without hint.

R4 — Integrated Retrieval

Can retrieve while solving a different or mixed problem.

These are proposed operational categories.

They are not yet validated psychometric levels.


Sensor 7 — Routing

Question

Can the learner identify which Mathematics is relevant when the method is not announced?

This sensor becomes increasingly important from P3 onward.

Blocked worksheet:

Percentage Practice

has already provided routing information.

Mixed assessment:

no topic label.

Now the child must discriminate.

We can test routing by placing several mathematical structures close together.

For example:

fraction

area

percentage

division

ratio

Then observe whether the learner identifies the relevant family before calculating.


Routing Error vs Knowledge Error

Suppose a learner selects division when multiplication is required.

Before reteaching multiplication, ask:

Does the child know multiplication?

If yes, the problem may be method selection.

This distinction is critical.

A student can possess the necessary Mathematics and still route incorrectly.

That requires a different intervention.


Sensor 8 — Transfer

Question

Does the learner retain control when the surface of the problem changes?

Transfer should be tested systematically.

A useful ladder is:

T0 — Exact Reproduction

Same structure and very similar surface.

T1 — Numerical Variation

Numbers change.

T2 — Language Variation

Wording changes.

T3 — Representation Variation

Diagram or mathematical form changes.

T4 — Context Variation

Story changes.

T5 — Integrated Transfer

Concept appears alongside another concept.

T6 — Unfamiliar Problem

No obvious stored template is available.

Again, these are proposed research levels rather than validated scales.

Their purpose is to give the research programme something explicit to test.


Why Transfer Matters

IES notes that Mathematics students often struggle to apply learning in new contexts and has funded research specifically examining which instructional conditions produce more fluent and transferable arithmetic learning. (ies.ed.gov)

So the State Estimator should not ask only:

Can the student repeat?

It should ask:

How far can the learning travel before it collapses?

That is a powerful measurement question.


Sensor 9 — Error Recurrence

Question

Does the same failure mechanism return?

One wrong answer may be noise.

Repeated related errors become a signal.

Examples:

  • repeated unit loss;
  • recurring denominator mistakes;
  • repeated wrong-operation selection;
  • repeated percentage reference errors;
  • repeated unfinished multi-step questions.

The useful measurement is not merely:

number of errors.

It is:

error genealogy.

Which errors are descendants of the same underlying mechanism?

This becomes central to PMRI-003.


Repair Should Change Error Recurrence

If an intervention was successful, we expect some future behaviour to change.

For example:

Before repair

same fraction-equivalence error repeatedly appears.

After repair

the error should:

  • disappear;
  • decrease substantially;
  • or change form in a way suggesting partial progress.

If the exact error continues unchanged, the intervention should be questioned.

This is the receiver test again.


Sensor 10 — Calibration and Self-Correction

Question

Can the learner detect when something is wrong without the tutor announcing it?

Examples:

A pencil is 350 metres long.

25% of a number is larger than 500% of the same positive number.

The area is written in centimetres rather than square centimetres.

A learner with stronger calibration may pause and investigate.

This is important because examinations cannot provide immediate corrective feedback.

The learner needs internal monitoring.


Self-Correction Is Strong Evidence

Consider two wrong responses.

Learner A

Produces wrong answer.

Accepts it immediately.

Learner B

Produces wrong answer.

Says:

“That doesn’t make sense.”

Returns to working.

Repairs it.

Both temporarily made an error.

Only one demonstrated effective error regulation.

So the State Estimator should record:

error recovery

not only error occurrence.


Sensor 11 — Performance Under Constraint

The preceding sensors examine capability.

Eventually we need to ask:

Does the capability survive load?

Possible constraints include:

  • mixed topics;
  • limited time;
  • no tutor prompts;
  • longer papers;
  • unfamiliar ordering;
  • examination pressure.

A learner can possess a capability in isolation but fail to dispatch it under load.

This is particularly important in P5, P6 and PSLE Mathematics.


Capability State vs Performance State

We therefore distinguish:

Capability State

What the learner can do under sufficiently supportive conditions.

Dispatchability State

How readily that capability can be brought online.

Performance State

How reliably it survives realistic constraints.

This prevents a common error:

assuming examination failure always means the Mathematics was never understood.

Sometimes it was understood but not reliably operational.


The State Estimator Matrix

The proposed observational matrix is:

SensorCore question
AccuracyWas the Mathematics correct?
LatencyHow readily did useful action begin?
ExplanationCan the relationship be communicated?
RepresentationCan the idea be transformed?
Prompt dependenceHow much help is required?
Delayed retrievalDoes the learning survive time?
RoutingCan the relevant Mathematics be selected?
TransferDoes learning survive surface change?
Error recurrenceDoes the same mechanism return?
CalibrationCan error be detected and repaired?
Constraint performanceDoes capability survive mixed/timed conditions?

The key point is:

No single column is Mathematics understanding.

Understanding is inferred from the pattern.


A Worked State-Estimator Example

Suppose a Primary 5 learner is reportedly:

“weak at percentage.”

We run several observations.

Observation 1

Direct percentage calculation:

80% of 250

Correct.

Observation 2

Student explains that percent means:

“out of one hundred.”

Reasonable conceptual evidence.

Observation 3

Can convert:

1/4 → 25%

Correct.

Observation 4

Word problem involving percentage increase.

Wrong.

Observation 5

Tutor says:

“What is the original 100%?”

Student immediately reconstructs the problem correctly.

Observation 6

New percentage increase problem without hint.

Wrong again.

What does this suggest?

Not necessarily:

percentage knowledge missing.

Instead, evidence may support:

routing/reference-quantity control is unstable and prompt dependence remains high.

So the intervention changes.

Less:

repeat direct percentage calculations.

More:

discriminate reference quantities across mixed percentage situations.

That is what state estimation is for.


Another Example: Same Score, Different Intervention

Two students score 6/10 on a fraction probe.

Student A

  • understands fraction diagrams;
  • explains equivalence;
  • retrieves slowly;
  • makes multiplication errors.

Likely intervention:

improve low-level retrieval/execution.

Student B

  • calculates familiar questions correctly;
  • cannot explain equivalence;
  • fails representation changes;
  • depends on memorised procedure.

Likely intervention:

rebuild conceptual/representation structure.

Same score:

6/10

Different estimated state.

Different repair.


Part 3 of 3

From Observation to Research Method

The State Estimator must not become tutor intuition dressed in technical language.

It needs a disciplined protocol.


Step 1 — Define the Target Capability

Do not assess:

Mathematics.

Assess something narrower.

For example:

equivalent fractions.

or:

multiplicative comparison.

or:

selecting between area and perimeter.

The more precise the target, the more interpretable the observations.


Step 2 — Establish a Baseline

Collect enough information to answer:

What can the learner currently do without intervention?

Use appropriately chosen tasks.

Record:

  • response;
  • method;
  • latency where meaningful;
  • representation;
  • prompts;
  • confidence;
  • self-correction.

This becomes the baseline state estimate.


Step 3 — Vary One Important Dimension

A good diagnostic probe changes something.

For example:

same concept, different number

or:

same concept, different representation

or:

same concept, different context

The change helps reveal what part of the previous success was robust.


Step 4 — Remove Support

If the student succeeds after tutoring:

reduce prompting.

This tests ownership.

Success under full scaffolding should not be reported as independent mastery.


Step 5 — Delay

Do not conclude repair from the immediate post-test alone.

Return after time.

This measures availability after recent instructional support has faded.


Step 6 — Mix

Place the concept among alternatives.

This tests routing.

The student now has to decide which mathematical system is relevant.


Step 7 — Transfer

Change the surface.

If the learning survives, confidence in the state estimate increases.


Step 8 — Apply Constraint

For older Primary students, test whether the capability survives:

  • time;
  • mixed papers;
  • longer working sequences.

This moves from learning measurement toward performance measurement.


Step 9 — Update the Hypothesis

The estimated state should change when evidence changes.

Suppose we initially believed:

concept weak.

After several probes we discover:

  • explanation strong;
  • representation strong;
  • untimed accuracy high;
  • timed performance poor.

The hypothesis should change toward:

execution/regulation bottleneck.

A research institution should reward corrected diagnosis.

Not defend its first guess.


The Intervention Test

One of the strongest ways to evaluate a diagnostic hypothesis is to intervene.

Suppose we believe:

multiplication retrieval is constraining fraction performance.

Prediction:

improving multiplication retrieval should reduce fraction-task latency or error under otherwise similar conditions.

If multiplication improves but fraction performance remains unchanged, the weak-link hypothesis receives less support.

Maybe the bottleneck lies elsewhere.

This turns intervention into a diagnostic experiment.


Diagnosis Should Make Predictions

A useful state estimate should predict something.

For example:

Hypothesis

Representation is the bottleneck.

Prediction

When supplied an appropriate diagram, performance should improve materially.

If not:

reconsider the hypothesis.

Or:

Hypothesis

Knowledge exists but retrieval is weak.

Prediction

Recognition will be substantially stronger than unaided delayed recall.

If recognition and recall are both absent:

the concept itself may be less established than expected.

Predictions make the system testable.


Competing Hypotheses

When evidence is ambiguous, keep multiple explanations alive.

Example:

Student fails two-step problems.

Candidate explanations:

H1 — sequencing weakness

H2 — multiplication retrieval bottleneck

H3 — translation weakness

H4 — regulation under load

Now design probes that distinguish them.

That is more scientific than immediately selecting whichever explanation sounds most plausible.


Falsification Condition

Each diagnosis should have a condition under which it weakens.

Example:

Proposed diagnosis

Routing weakness.

Weakening evidence

Student selects appropriate methods reliably across several mixed tasks without prompting.

Then routing is probably not the main bottleneck.

Move elsewhere.

This prevents categories from becoming unfalsifiable explanations for everything.


No Universal State Score Yet

It may be tempting to take the eleven sensors and produce:

Mathematics State = 82.7

We should not do this yet.

There are several unanswered questions:

  • Are the sensors independent?
  • Which are task-specific?
  • Which are age-specific?
  • How reliable are tutor judgements?
  • How should latency be normalised?
  • Is transfer one dimension or several?
  • How should prompt dependence be weighted?
  • Does one combined score improve intervention?

Until those questions are studied:

preserve the profile.

Do not compress prematurely.


The Mathematics State Profile

A future parent-facing report could look conceptually like:

Fraction Concept

Stable

Retrieval

Moderately stable

Representation

Stable

Routing

Needs support

Transfer

Unstable

Prompt Dependence

Moderate

Error Recurrence

High in mixed problems

Suggested Next Intervention

Mixed discrimination + delayed transfer retest

This is more useful than inventing:

Fraction Intelligence Score: 67.


Observation Reliability

A research institution also needs to ask:

Would two trained tutors interpret the same performance similarly?

This is an important future validation requirement.

The State Estimator cannot depend entirely on one tutor’s subjective impression.

Potential research steps include:

  • explicit observation definitions;
  • example responses;
  • coding rubrics;
  • double coding;
  • agreement checks;
  • adjudication of ambiguous cases.

This will become necessary before stronger scientific claims can be made.


Measurement Validity

Another question:

Are we measuring what we think we are measuring?

Suppose latency is used as a retrieval measure.

But a child is naturally deliberate and slow across all tasks.

Then latency may partly capture response style rather than mathematical retrieval.

So each sensor needs validity testing.

This is exactly why the present framework is labelled:

MSE v0.1

It is a proposed measurement system.

Not a finished one.


State Estimator Research Agenda

The first validation programme should investigate:

RQ1

How reliably can tutors code the proposed observations?

RQ2

Which sensors provide information beyond score alone?

RQ3

Which sensor combinations best predict independent performance?

RQ4

Which observations predict delayed retention?

RQ5

Which observations predict transfer?

RQ6

Can prompt dependence predict later independent performance?

RQ7

Can error recurrence distinguish weak-link types?

RQ8

Does routing performance on mixed tasks predict examination performance better than blocked-practice accuracy?

RQ9

How stable are state estimates across days?

RQ10

Which state dimensions change most after targeted intervention?

These are empirical questions.

They form part of the institution’s future research programme.


Singapore Alignment Without Overclaiming

SEAB’s adaptive Mathematics tools provide an important external precedent for going beyond a single total score.

MathsCheckPlus is designed to identify readiness and support targeted remediation, while CATalytics provides more focused diagnostic information on prior knowledge and learning gaps. (SEAB)

CATalytics currently assesses P5/P6 proficiency in selected Standard Mathematics topics and supplies student profiles, descriptors and common-error information. (SEAB)

eduKateSG should be precise about the relationship.

We are not claiming:

eduKateSG’s State Estimator is the same as SEAB’s assessment system.

Nor:

SEAB validates our specific categories.

What SEAB does validate is the importance of diagnostic resolution:

understanding where the learning gap lies can support more targeted intervention than score alone.

Our research question begins from there.


What a Small-Group Environment Adds

Standardised assessments have enormous strengths:

  • consistency;
  • coverage;
  • comparability;
  • scalable reporting.

A three-student tuition environment offers different potential information.

The tutor can observe:

  • strategy before the answer;
  • hesitation;
  • self-talk;
  • representation choice;
  • response to prompts;
  • self-correction;
  • what happens when support is removed.

So the two environments need not compete.

They answer different measurement questions.

The research opportunity is:

Can high-resolution observational evidence improve the instructional usefulness of formal assessment evidence?

That is a much more defensible institutional research question.


The Lesson as a Repeated Measurement Environment

A normal lesson contains many natural probes.

Student attempts independently.

Observation.

Tutor asks a strategic question.

Observation.

Student changes representation.

Observation.

Hint is removed.

Observation.

Topic returns next week.

Observation.

Same Mathematics appears in a mixed paper.

Observation.

This means measurement can occur inside teaching without requiring constant formal testing.

But only if the observations are recorded systematically enough to become meaningful.


Do Not Turn the Child Into a Dashboard

This is an important boundary.

More measurement is not always better.

A learner is not:

a collection of performance indicators.

The purpose of the State Estimator is to make intervention more precise while keeping teaching human, calm and proportionate.

So:

Measurement serves learning.

Learning does not serve measurement.


Parent Communication

A research-oriented tuition organisation should also improve how findings are communicated to parents.

Instead of:

“Your child is weak in Maths.”

report:

“Whole-number calculation is generally stable. Fraction concepts are understood when represented visually, but delayed retrieval and mixed-question selection remain inconsistent. We are therefore working on retrieval and routing before increasing paper volume.”

This is:

  • more precise;
  • more actionable;
  • less identity-forming.

It also tells the parent what is being tested next.


Student Communication

For the learner:

Instead of:

“You don’t know fractions.”

try:

“You understand the fraction when we draw it. The part we are strengthening now is recognising when to use that idea without a hint.”

This changes the psychological meaning of diagnosis.

The weakness becomes:

a specific trainable state.

Not:

a verdict about mathematical ability.


Tutor Communication

The tutor’s internal record should answer:

What was observed?

What hypothesis does it support?

How confident are we?

What alternative explanation remains?

What intervention follows?

What result would confirm or weaken the hypothesis?

This is the beginning of research discipline inside ordinary practice.


The MSE v0.1 Observation Record

A minimal future record could contain:

Target

What Mathematics is being observed?

Task

What did the learner attempt?

Accuracy

What happened?

First Divergence

Where did the reasoning first leave the correct path?

Latency

Was response unusually immediate, normal or delayed?

Representation

What representation was selected?

Prompt Level

How much support was required?

Explanation

What relationship could the learner articulate?

Retrieval

Was the Mathematics independently available?

Transfer

Did it survive variation?

Self-Correction

Did the learner detect error?

Current Hypothesis

What mechanism currently best explains the pattern?

Confidence

Low / Moderate / Higher

Competing Hypothesis

What alternative still fits?

Intervention

What will we change?

Retest

What result will tell us whether the intervention worked?

This is enough structure to make the process inspectable without turning every lesson into laboratory bureaucracy.


Evidence Boundary

What Current Evidence Supports

Current assessment and education research strongly supports using assessment evidence to inform instruction rather than treating assessment only as an endpoint. SEAB’s adaptive Mathematics tools explicitly identify learning gaps and provide targeted information for remediation. (SEAB)

Research also treats mathematical fluency and transfer as distinguishable outcomes and continues to investigate instructional conditions that help learners apply arithmetic knowledge beyond trained problems. (ies.ed.gov)


What PMRI-002 Adds

eduKateSG proposes a higher-resolution observational framework using:

accuracy

latency

explanation

representation

prompt dependence

delayed retrieval

routing

transfer

error recurrence

calibration

performance under constraint

These dimensions are research constructs requiring validation.


What We Do Not Yet Know

We do not yet know:

  • which observations are redundant;
  • which observations are most predictive;
  • how reliable tutor coding will be;
  • whether eleven sensors are too many;
  • whether some sensors should split;
  • whether some should merge;
  • which measures work best at different Primary levels;
  • whether the resulting state estimates improve learning outcomes compared with simpler diagnostic approaches.

Those questions are not weaknesses in the research programme.

They are the research programme.


The eduKateSG Mathematics State Estimator

The complete MSE v0.1 runtime is:

Define Target

What mathematical capability are we investigating?

Observe

What does the learner actually do?

Decompose

Where does performance first diverge?

Sample Multiple Sensors

Accuracy
Latency
Explanation
Representation
Prompt Dependence
Retrieval
Routing
Transfer
Error Recurrence
Calibration
Constraint Performance

Triangulate

Which observations converge?

Generate Hypotheses

What mechanisms could explain the pattern?

Assign Confidence

Low / Moderate / Higher

Keep Alternatives

What else could explain it?

Intervene

Change one useful variable.

Immediate Retest

Did behaviour change?

Remove Support

Does success remain?

Delay

Does the learning return later?

Mix

Can the learner route correctly?

Transfer

Does learning survive surface change?

Constraint Test

Does capability remain available under load?

Update State

What does the new evidence support?

Record

What did we learn?

Research Memory

Does the case support or challenge the framework?


From “How Many Marks?” to “What State Produced the Marks?”

Marks remain important.

The State Estimator does not replace them.

It gives us a way to ask the next question.

A student scores:

68%

We ask:

What mathematical state generated the 68%?

Another student scores:

68%

We ask again.

The answer may be different.

That difference is where targeted education begins.


Conclusion

Understanding Must Be Inferred Carefully

There is no single worksheet question that reveals everything a child understands.

There is no single mark that fully describes the learner.

There is no single wrong answer that identifies one unique cause.

So the research problem is:

How do we infer the learner’s state without pretending that our inference is the learner?

The eduKateSG Mathematics State Estimator begins with a simple discipline:

observe first

infer cautiously

triangulate

test the inference

change the intervention

observe again

That is measurement as a learning loop.

Not measurement as labelling.

The child may:

recognise but not retrieve

retrieve but not route

route but execute poorly

execute correctly but fail transfer

transfer without time pressure but collapse during examination conditions

Each state produces a different educational problem.

And therefore potentially a different repair.

That is why:

Score ≠ State

and:

State ≠ Identity

The goal is not to construct the most complicated diagnostic system possible.

It is to obtain just enough resolution to make the next educational decision better.

When the evidence is weak:

remain uncertain.

When several observations converge:

form a stronger hypothesis.

When the intervention fails:

revise the diagnosis.

When transfer fails:

do not declare mastery.

When delayed retrieval fails:

do not confuse immediate performance with durable learning.

When support is still required:

do not report independent control.

And when the learner changes:

update the state.

That is the Mathematics State Estimator.

Not a final test.

A continuously improving estimate of:

What can this learner currently do, under what conditions, with how much support, and what needs to change next?

That is the measurement architecture needed before eduKateSG can credibly build the next research layer.


Research Status at Publication

PMRI-001

Primary Mathematics Learning-System Framework established.

PMRI-002

Mathematics State Estimator v0.1 established as a proposed measurement framework.

Next Research Paper:

PMRI-003 — The Primary Mathematics Weak-Link Atlas

Why the Same Wrong Answer Can Have Different Causes

Primary research question:

Can recurring mathematical failure patterns be classified into a small enough set of diagnostic categories to improve intervention selection without oversimplifying the learner?

Candidate failure classes:

Missing Node

Broken Edge

Weak Link

Wrong Edge

Routing

Translation

Transfer

Calibration

Regulation

PMRI-003 will attempt to convert those categories from useful language into a falsifiable diagnostic taxonomy.