VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

eduKateSG Primary Mathematics Research Agenda 2026–2030

What We Know, What We Think, What We Need to Test

PMRI-006 — Primary Mathematics Research Institute Series
eduKateSG Learning Sciences & Education Research Programme
Framework: PMRI Research Charter RC v1.0
Status: Founding Research Agenda / Institutional Research Charter
Research Horizon: 2026–2030
Primary Domain: Primary Mathematics / Learning Sciences / Diagnostic Education / Educational Technology
Jurisdictional Context: Singapore Primary 1–6 / PSLE Mathematics
Publication Date: August 2026


Research Classification

Article Type: Institutional research agenda + methods charter + governance framework

Purpose:

This paper defines how eduKateSG intends to move from:

research-informed teaching

towards:

a cumulative programme of inspectable education research

focused initially on Primary Mathematics.

It establishes:

  • the research questions;
  • research programmes;
  • evidence hierarchy;
  • methods ladder;
  • replication policy;
  • publication standards;
  • null-result policy;
  • version-control system;
  • student-data boundary;
  • AI research boundary;
  • research-to-practice loop;
  • and the milestones required before stronger scientific claims are made.

Quick Read

eduKateSG should not become a research institution because we write:

“eduKateSG is a research institution.”

It should become research-like because readers can inspect:

What question did we ask?

What evidence did we use?

What did we predict?

How did we test it?

What happened?

What failed?

What changed in our model?

Can somebody else repeat the work?

That is the institutional shift.

The first five PMRI papers established:

PMRI-001 — System

What is the Primary Mathematics learning system?

PMRI-002 — Measurement

How do we estimate learner state?

PMRI-003 — Failure

Where and how does the system break?

PMRI-004 — Continuity

How does earlier Mathematics survive inside later Mathematics?

PMRI-005 — Repair

How do we know whether intervention actually worked?

PMRI-006 adds the final layer:

How will eduKateSG test, revise and publish these ideas over the next four years?

The research loop becomes:

Question

Evidence

Hypothesis

Protocol

Observation / Experiment

Result

Replication

Failure Analysis

Model Revision

Teaching Change

New Question

The purpose is not to prove eduKateSG right.

The purpose is to build a system that becomes less wrong over time.


The Institutional Principle

NIE’s 2026 discussion of education research for impact describes useful education research as needing to be rigorous, relevant and responsive, and places particular emphasis on research that can shape classrooms, policy and learner experiences. (singteach.nie.edu.sg)

That is an excellent institutional standard for eduKateSG.

But we add a fourth requirement:

Inspectable.

Our research should aim to be:

Rigorous

Can the evidence support the claim?

Relevant

Does the question matter to learning?

Responsive

Does the research change when learner needs, evidence or educational conditions change?

Inspectable

Can an external reader understand how we reached the conclusion?

This becomes:

The eduKateSG 4R Standard

Rigorous → Relevant → Responsive → Reviewable


Research Institution vs Research Marketing

There is an important boundary.

A research-marketing organisation says:

“Research proves our method works.”

A research organisation asks:

Which research?

Which learners?

Which outcome?

Which comparison?

Over what period?

Did transfer occur?

Did the effect replicate?

What did not improve?

What remains uncertain?

That difference will govern PMRI publications.


The Strong Claim Rule

The stronger the claim:

the stronger the evidence required.

A classroom observation may justify:

“We repeatedly observe this pattern.”

It does not automatically justify:

“This causes improvement in Primary Mathematics students.”

A controlled study may justify more.

A replicated controlled study may justify more again.

eduKateSG will therefore distinguish claim strength from confidence enthusiasm.

We do not strengthen claims because we like the intervention.

We strengthen them because the evidence warrants it.


Part 1 of 3

Building the eduKateSG Research Institution

Research Standard 1 — Every Programme Begins With a Question

No research article should begin with:

Here is the answer.

It should begin with:

What are we trying to understand?

Examples:

Does mixed practice reveal routing weaknesses that blocked practice hides?

Does successful fraction repair persist four weeks later?

Which Primary 2 mathematical states best predict difficulty in Primary 3?

Does reducing tutor prompting increase independent transfer?

Can an AI-assisted error classifier improve tutor diagnostic consistency without replacing tutor judgement?

Questions come before conclusions.


Research Standard 2 — Every Major Claim Receives an Evidence Class

The PMRI evidence system remains:

E1 — Established External Evidence

Supported by sufficiently strong external evidence or authoritative sources.

E2 — eduKateSG Evidence Synthesis

A reasoned synthesis across research sources.

E3 — Structured Practice Observation

A recurring eduKateSG pattern recorded systematically but without strong causal inference.

E4 — Working Hypothesis

A mechanism proposed for testing.

E5 — Design Principle

An intervention rule used operationally while remaining open to revision.

But PMRI-006 adds experimental claim levels.


Study-Evidence Levels

S0 — External Evidence Review

No eduKateSG learner data.

Purpose:

understand existing research.


S1 — Structured Observation

Teaching behaviour is systematically described.

Purpose:

generate hypotheses.

Cannot by itself establish causality.


S2 — Repeated-Measure / Single-Learner Study

Repeated observations before and after intervention.

Purpose:

investigate whether behaviour changes in a particular learner or small set of learners.

Useful for mechanisms.

Limited generalisability.


S3 — Cohort Study

Patterns are examined across a larger group of learners.

Purpose:

estimate associations and common trajectories.

Still limited for causal claims.


S4 — Quasi-Experimental Study

A comparison is constructed without full randomisation.

Purpose:

strengthen causal inference where appropriate.


S5 — Randomised Study

Where educationally justified, ethically appropriate and practically feasible, learners or interventions are randomly assigned.

Purpose:

stronger causal estimation.

Not every useful education question requires this design.


S6 — Independent or Multi-Site Replication

The finding is tested again:

  • with new learners;
  • by other tutors;
  • in another location;
  • or ideally by an external group.

Purpose:

determine whether the result survives beyond the original context.


Why We Need a Methods Ladder

Not every research question should be forced into one method.

If we ask:

What errors repeatedly occur during P3 multiplication?

structured observation may be appropriate.

If we ask:

Does Intervention A outperform Intervention B?

a stronger comparative design may be required.

The method should match the question.

IES’s Standards for Excellence in Education Research similarly distinguish the importance of rigorous theory, measurement, implementation, replication and transparent methods, rather than treating one research design as sufficient for every problem. (ies.ed.gov)


Research Standard 3 — Separate Exploration From Confirmation

Early work often asks:

What might be happening?

This is exploratory research.

Later work asks:

Does our previously specified prediction survive a formal test?

This is confirmatory research.

The two should not be silently mixed.


Exploratory Research

Useful for:

  • discovering error patterns;
  • generating weak-link categories;
  • identifying candidate dependencies;
  • finding unexpected student strategies.

Exploration is valuable.

It should be labelled.


Confirmatory Research

Before collecting the decisive data, specify:

  • research question;
  • primary hypothesis;
  • primary outcome;
  • sample;
  • analysis;
  • exclusion rules;
  • stopping rule where relevant.

IES explicitly encourages preregistration as part of its open-science standards. (ies.ed.gov)

eduKateSG should progressively adopt the same principle for studies making stronger confirmatory claims.


The eduKateSG Research Register

Beginning with PMRI implementation, each formal study should receive an identifier.

For example:

PMRI-STUDY-2027-001

The public research register should record:

Study Title

Research Question

Study Classification

S1 / S2 / S3 / S4 / S5 / S6

Status

Proposed / Preregistered / Active / Completed / Replication / Closed

Hypothesis

Where applicable.

Primary Measures

Declared before confirmatory analysis.

Protocol Version

Amendments

Every meaningful change logged.

Results

Positive / mixed / null / contradictory.

Publication

Where the complete report can be read.

This gives the research programme memory.


Research Standard 4 — Publish the Methods

A research result is difficult to evaluate without knowing how it was produced.

At minimum, formal PMRI studies should eventually disclose:

  • participant characteristics at an appropriate, privacy-preserving level;
  • mathematical target;
  • tasks used;
  • observation definitions;
  • intervention;
  • duration;
  • prompt conditions;
  • outcome definitions;
  • analysis procedures;
  • exclusions;
  • missing-data treatment;
  • limitations.

IES’s SEER framework similarly emphasises open methods, detailed descriptions of measurement and analysis, and transparency around research activity. (ies.ed.gov)


Research Standard 5 — Publish Null Results

This is essential.

Suppose eduKateSG predicts:

Interleaving will improve transfer.

The study finds:

no detectable advantage in this implementation.

Publish it.

Suppose:

early fraction repair should improve later percentage performance.

No spillover appears.

Publish it.

Suppose the Weak-Link Atlas category:

Broken Edge

cannot be coded reliably.

Publish it.

A null result is not a failed publication.

It is information.


The Null-Results Ledger

PMRI should maintain a public:

Null / Contradictory Findings Register

Each entry records:

Prediction

Observed Result

Interpretation

Model Consequence

Examples:

Finding NR-001

Proposed dependency did not produce measurable downstream change.

Action: edge downgraded from Strong Dependency to Conditional Dependency.

Finding NR-002

Tutor inter-rater agreement on Wrong Edge classification was poor.

Action: category definition revised.

Finding NR-003

Intervention improved immediate performance but not four-week retrieval.

Action: repair status limited to immediate acquisition.

This would be one of the clearest signals that eduKateSG is behaving as a research organisation rather than a marketing organisation.


Research Standard 6 — Replication Before Institutionalisation

One result should rarely become:

“The eduKateSG Method.”

A new finding should pass several gates.

Initial Finding

Repeat

New Cohort

Different Tutor

Different Topic or Level where relevant

External Replication where possible

Only then should confidence rise substantially.

Replication is a major concern across education research; IES initiatives explicitly identify scarce replication as a problem and are developing research infrastructure to support larger, more representative and replicable studies. (ies.ed.gov)


The Replication Ladder

R1 — Same Tutor, New Learners

R2 — Different eduKateSG Tutor

R3 — Different eduKateSG Location

R4 — Different Primary Level

where theory predicts generalisation.

R5 — External Collaborator

R6 — Independent Replication

Confidence rises with appropriate successful replication.

Failure to replicate triggers model review.


Research Standard 7 — Version Everything Important

The existing frameworks already begin this process.

PM-LS v1.0

MSE v0.1

WLA v1.0

MCM v1.0

MRVP v1.0

PMRI-006 formalises it.

Every model should have:

Version

Date

Change

Evidence Trigger

Consequence


Example Model Change

WLA v1.0

Contains:

Missing Node / Broken Edge / Weak Link / Wrong Edge / Routing / Translation / Transfer / Calibration / Regulation.

Suppose research shows tutors cannot reliably distinguish Broken Edge from Wrong Edge.

Then:

WLA v1.1

might merge or redefine them.

The public changelog should state:

changed because inter-rater reliability did not meet the predeclared threshold.

That is cumulative science.


Research Standard 8 — Protect the Learner Before Protecting the Research

The research programme exists because children are learning.

Research cannot be allowed to damage that purpose.

Therefore:

Learning welfare outranks experimental convenience.

No learner should be denied clearly necessary instruction simply to preserve a clean research comparison.

No intervention should be continued when evidence suggests it is educationally harmful.

No data should be collected merely because it might someday be interesting.


Children’s Data Requires a Higher Standard

Singapore’s PDPC has specific guidance for children’s personal data and explicitly includes technology-aided learning within the relevant digital environment. The guidance emphasises appropriate notification and consent, data protection by design and data minimisation; PDPC also recommends Data Protection Impact Assessments for services involving children’s data. (PDPC)

For PMRI, this produces a strict default:

Collect the minimum learner data necessary to answer the research question.


The PMRI Data-Minimisation Rule

Before collecting any field, ask:

Why do we need this?

If there is no research or operational reason:

do not collect it.

Potential research fields should be traceable to:

  • a stated purpose;
  • an authorised user;
  • a retention rule;
  • a protection rule.

Separate Operations Data From Research Data

A tuition organisation naturally holds operational information.

Research should not automatically consume all of it.

The architecture should distinguish:

Operational Dataset

Needed to operate lessons and communicate with families.

Research Dataset

Contains only information explicitly justified by the research protocol.

Publication Dataset

De-identified or aggregated to the extent necessary so published research does not expose individual learners.

This separation reduces risk.


Longitudinal Research Needs Special Discipline

The PMRI programme is particularly interested in:

P1 → P2 → P3 → P4 → P5 → P6 → PSLE

Longitudinal data is scientifically valuable.

It also increases privacy responsibility.

Current PDPC guidance stresses data minimisation and not retaining personal data once the purpose no longer requires it; where longer-term analytics genuinely requires retention, anonymised or aggregated forms should be considered where appropriate. (PDPC)

Therefore:

longitudinal value does not justify indefinite raw-data retention.


Research Standard 9 — Research Should Return to Teaching

NIE’s SingTeach has operated specifically around bridging education research and classroom practice, and its 2026 research-impact work again emphasises connecting research with educator practice. (singteach.nie.edu.sg)

PMRI adopts the same general objective.

Research should not terminate at:

PDF published.

The loop should continue.

Finding

Teaching Design

Tutor Training

Classroom Implementation

Observation

New Evidence

This is the eduKateSG research-to-practice loop.


Research Standard 10 — Practice Should Generate New Research Questions

Knowledge also flows in the opposite direction.

Tutors may repeatedly observe:

students know fractions but cannot recognise them in ratio contexts.

That becomes a research question.

Or:

one visual representation appears to reduce prompting dramatically.

Research question.

Or:

interleaving helps one learner state and overwhelms another.

Research question.

The classroom is therefore not merely where research is implemented.

It is one source of new scientific problems.


Part 2 of 3

The eduKateSG Primary Mathematics Research Programmes 2026–2030

The founding programme will contain ten research lanes.

Together they test the architecture created in PMRI-001 to PMRI-005.


Research Lane 1 — Learner State Estimation

Core Question

Can we estimate a learner’s Mathematics state more usefully than score alone?

Primary constructs:

  • accuracy;
  • latency;
  • explanation;
  • representation;
  • prompt dependence;
  • retrieval;
  • routing;
  • transfer;
  • calibration;
  • regulation.

Initial Studies

MSE-01

Inter-rater reliability of tutor coding.

MSE-02

Which sensors predict independent performance?

MSE-03

Which sensors predict delayed retrieval?

MSE-04

Does the MSE profile provide intervention information beyond test score?


Success Condition

The State Estimator should remain only if it improves educational decisions.

If:

Score + ordinary tutor judgement

performs equally well,

then the additional measurement architecture needs justification or simplification.


Research Lane 2 — Weak-Link Diagnosis

Core Question

Are the nine Weak-Link Atlas categories distinguishable and intervention-useful?

Candidate classes:

  • Missing Node;
  • Broken Edge;
  • Weak Link;
  • Wrong Edge;
  • Routing;
  • Translation;
  • Transfer;
  • Calibration;
  • Regulation.

Research Sequence

Coding Reliability

Prediction

Intervention Match

Outcome

The strongest test is not:

Can tutors give errors different names?

It is:

Does the category help us choose an intervention that works better?


Research Lane 3 — Mathematics Learning Continuity

Core Question

Which earlier mathematical capabilities meaningfully influence later learning?

The goal is to build:

probabilistic dependency maps

rather than one rigid P1→P6 chain.


Initial Topic Maps

Multiplication Continuity Map

Fraction Continuity Map

Decimal–Percentage Continuity Map

Problem-Solving Routing Map

Primary-to-Secondary Mathematics Transition Map

Each should be tested locally before joining the global Mathematics map.


Research Lane 4 — Repair Validation

Core Question

When is a weakness truly repaired?

Candidate sequence:

Immediate

Independent

Delayed

Representation Change

Mixed Routing

Transfer

Constraint Performance

Research will test which stages predict future success.


Repair Durability Programme

Possible measures:

Recurrence Rate

How often does the original weakness return?

Reactivation Cost

How much support is required after forgetting?

Transfer Radius

How far does the capability generalise?

Prompt Independence

How much scaffolding remains necessary?

These are currently conceptual research constructs.

They should become formal metrics only if they can be operationalised reliably.


Research Lane 5 — Retrieval, Spacing and Interleaving

Core Question

Which practice architecture works best for which mathematical state?

This lane should explicitly reject:

one technique for everyone.

Instead test interactions.

For example:

novice concept + blocked practice

versus:

stable concept + interleaving

or:

short delay

versus:

longer delay

for different types of Mathematics.

The PMRI programme should treat learning strategies as conditional interventions.


Research Lane 6 — Mathematical Representation

Core Question

Which representations make particular mathematical structures easier to perceive, retrieve and transfer?

Possible research areas:

  • bar models;
  • number lines;
  • physical manipulatives;
  • diagrams;
  • tables;
  • symbolic equations;
  • dynamic digital representations.

Questions include:

When should representation be supplied?

When should the learner construct it?

When should representation be removed?

Does representation help transfer or produce new dependence?

This links directly to the Representation State in PMRI-001 and PMRI-002.


Research Lane 7 — Tutor Prompting and Independence

This could become one of eduKateSG’s most distinctive programmes because very small groups allow high-resolution observation of tutor–student interaction.

Core Question

How much help produces learning, and when does help begin producing dependence?


Prompt-Ladder Research

Possible levels:

P0 — Demonstration

P1 — Strategic Direction

P2 — Guiding Question

P3 — Minimal Cue

P4 — Independent

Study:

  • how quickly students move down the prompt ladder;
  • whether low prompt dependence predicts transfer;
  • whether some students are released too quickly;
  • whether others are over-supported.

The Tutor-Paradox Question

Can better tutoring sometimes look like less tutoring because the learner is doing more of the cognitive work?

This is a strong research question.

Tutor activity is not identical to learner learning.


Research Lane 8 — Examination Performance Conversion

Core Question

Why does available Mathematical capability fail to become examination marks?

This programme begins particularly in P5, P6 and PSLE.

Candidate loss classes:

  • knowledge;
  • retrieval;
  • translation;
  • routing;
  • sequencing;
  • calculation;
  • recording;
  • calibration;
  • timing;
  • recovery.

Capability-to-Performance Gap

Estimate:

What learner can do under sufficiently supportive untimed conditions

versus:

What learner produces under examination constraints

The difference becomes a research surface.


Examination Research Questions

Which losses are most recoverable?

Which checking protocols reduce recurrence?

Does targeted timing training improve score without damaging accuracy?

How does one difficult question propagate into later errors?

Can recovery training contain local failure?

For high-performing students, what predicts variance?

This creates the future Examinations and Marks Strategy Research Programme.


Research Lane 9 — Small-Group Learning

eduKateSG operates very small teaching groups.

That creates a natural research question.

Not:

Are three students automatically better than ten?

That would be simplistic.

Instead:

What becomes observable and educationally actionable when instructional group size is very small?

Potential variables:

  • tutor response latency;
  • amount of individual mathematical explanation;
  • number of observable strategies;
  • prompting frequency;
  • correction time;
  • independent work time;
  • peer explanation;
  • transfer opportunities.

This turns “small group” from a marketing adjective into a research programme.


Research Lane 10 — AI-Assisted Mathematics Research

This is the frontier lane.

MOE’s current EdTech Masterplan explicitly includes AI-supported teaching and learning, and Singapore’s SLS already uses AI-enabled features designed with pedagogical guardrails. (Ministry of Education)

eduKateSG should investigate AI.

But under strict boundaries.


AI Research Question 1

Can AI classify mathematical error patterns consistently enough to assist tutors?

Not:

replace tutor diagnosis.

Initial test:

Tutor Coding

versus:

AI Coding

versus:

Tutor + AI

Which is:

  • more accurate;
  • more consistent;
  • faster;
  • more useful?

AI Research Question 2

Can AI identify possible weak-link hypotheses from student working?

AI might say:

Candidate explanations: routing / fraction equivalence / arithmetic retrieval.

The human decides:

what to investigate.

AI assists hypothesis generation.

It does not make the learner-state verdict.


AI Research Question 3

Can AI generate discriminating diagnostic probes?

Given competing hypotheses:

H1 — Translation

H2 — Routing

AI could propose questions designed to distinguish them.

These questions would still require human review before use.


AI Research Question 4

Can AI assist longitudinal pattern detection?

Example:

Across months of de-identified error records:

does a recurring multiplication pattern precede later fraction difficulty?

AI may help identify candidate patterns at scales difficult for tutors to inspect manually.

Those patterns become:

hypotheses.

Not automatic truths.


The AI Human-Accountability Rule

Singapore’s 2026 Model AI Governance Framework for Agentic AI stresses meaningful human accountability, risk bounding, technical controls and appropriate human checkpoints; it explicitly states that humans remain ultimately accountable. (IMDA)

PMRI will adopt an even stricter educational rule:

AI may recommend. Humans remain responsible.

For high-impact decisions affecting a child:

  • level placement;
  • major intervention;
  • publication interpretation;
  • learner labelling;

AI output should not become the sole decision mechanism.


AI Confidence ≠ Scientific Confidence

A language model can generate a confident explanation.

That does not mean the explanation is true.

Therefore AI-generated analysis requires:

Source

Verification

Human Evaluation

Evidence Classification

AI fluency must never be mistaken for empirical support.


AI Research Must Be Audited

Research questions include:

  • Does the AI reproduce tutor biases?
  • Does it overdiagnose?
  • Does it produce different labels for equivalent student work?
  • Does model version change alter results?
  • Can the output be reproduced?
  • Does AI assistance improve or reduce tutor diagnostic quality?

AI itself becomes an object of research.

Not merely a research tool.


AI Data Boundary

If learner data is processed through AI systems, PMRI must apply the same or stronger privacy discipline.

PDPC’s current AI guidance addresses the responsibilities of organisations using personal data for AI system development and deployment, while Singapore’s broader AI governance framework emphasises transparency, accountability and human-centred oversight. (PDPC)

So PMRI’s default will be:

No identifiable learner data should be sent into an external AI system merely because doing so is convenient.

Use:

  • minimised data;
  • de-identification;
  • controlled systems;
  • documented purpose;
  • appropriate permissions;

where AI-assisted research is genuinely justified.


Part 3 of 3

The 2026–2030 Research Roadmap

The programme should grow in phases.

Do not attempt everything at once.


Phase 0 — 2026

Build the Research Infrastructure

2026 is the foundation year.

Primary objectives:

Publish PMRI-001 to PMRI-006

Establish the intellectual architecture.

Create the Research Register

Every formal study gets an ID.

Create the Evidence Labels

E1–E5 / S0–S6.

Create the Model Registry

PM-LS / MSE / WLA / MCM / MRVP.

Create the Changelog

Track every model revision.

Create the Null-Results Ledger

Record contradictions.

Create the Data-Governance Protocol

Separate operational and research information.

Create Observation Rubrics

Begin making tutor observations codeable.

The objective of 2026 is not:

prove every model.

It is:

make future testing possible.


2026 Output

By the end of the initial infrastructure phase, a reader should be able to visit eduKateSG and find:

Research

Research Programmes

Research Questions

Frameworks

Methods

Current Studies

Results

Null Results

Model Changes

Research-to-Practice Articles

That changes how the entire organisation reads.


Phase 1 — 2027

Measurement and Reliability

Before testing ambitious interventions, determine whether the instruments are sufficiently reliable.

Priority:

Mathematics State Estimator

Can tutors apply it consistently?

Weak-Link Atlas

Can categories be distinguished?

Prompt Levels

Can support be coded reliably?

Transfer Ladder

Can transfer levels be operationalised?

Repair Validation

Can repair states be identified consistently?

This is measurement infrastructure.


2027 Core Studies

STUDY-MSE-01

Tutor inter-rater agreement.

STUDY-WLA-01

Weak-link category discrimination.

STUDY-MRVP-01

Immediate vs delayed repair classification.

STUDY-PRMPT-01

Prompt coding and independent performance.

STUDY-TRNS-01

Near vs representation vs contextual transfer.

The output should be:

fewer but more defensible constructs.


Phase 2 — 2028

Learning Mechanisms

Once measurement improves, investigate mechanisms.

Priority programmes:

  • retrieval;
  • spacing;
  • interleaving;
  • representation;
  • prompt fading;
  • weak-link repair;
  • continuity.

Questions become:

Which interventions change which learner states?


The Interaction Principle

Avoid asking only:

Does interleaving work?

Ask:

For whom, when, and after what prerequisite state?

Avoid:

Does a bar model work?

Ask:

For which mathematical structures and which learner states does the representation improve reasoning or transfer?

Avoid:

Does retrieval work?

Ask:

What kind of retrieval, for what Mathematics, over what delay?

This creates a conditional science of Mathematics learning.


Phase 3 — 2029

Longitudinal and Predictive Research

By 2029, enough structured information may exist to examine:

P1 → P2

P2 → P3

P3 → P4

P4 → P5

P5 → P6

P6 → PSLE

Research questions include:

Which early states predict later repair demand?

Which weaknesses disappear naturally?

Which propagate?

Which repairs persist for years?

Which students follow alternative successful trajectories?

This is where the Mathematics Continuity Map becomes genuinely longitudinal.


Prediction Requires Restraint

A predictive model should not become a deterministic learner label.

If an early state predicts elevated risk:

use it to create support.

Do not turn:

increased probability

into:

fixed future.

The system should remain adaptive.


Phase 4 — 2030

Replication, Generalisation and Institutional Review

By 2030, the programme should ask:

Which findings survived?

Not:

How many articles did we publish?

Assess:

Replication

Which results repeated?

Generalisation

Which worked across tutors, cohorts and locations?

Failure

Which original models were wrong?

Simplification

Which framework components can be removed?

External Collaboration

Which findings can be tested outside eduKateSG?

Research Impact

Which research actually changed teaching?

This is the first major programme review.


The 2030 Test

The programme has succeeded if eduKateSG can say:

“Here are the models we started with in 2026.”

“Here are the ones that survived.”

“Here are the ones that failed.”

“Here is why they failed.”

“Here is the evidence that changed our teaching.”

That is more scientifically meaningful than:

“We published 500 research articles.”


Publication Architecture

The PMRI website should eventually expose six layers.

Layer 1 — Research Agenda

What are we investigating?

Layer 2 — Frameworks

What models are currently being tested?

Layer 3 — Studies

What tests have been run?

Layer 4 — Results

What happened?

Layer 5 — Model Changelog

What changed?

Layer 6 — Practice Translation

What changed in tuition?

This lets parents, researchers, tutors and AI systems enter at different depths.


Parent Layer

Plain language:

What does this mean for my child?


Practitioner Layer

Operational:

How should teaching change?


Research Layer

Technical:

What was measured, how and with what limitations?


Machine-Readable Layer

Structured fields such as:

Study ID

Research Question

Population

Method

Outcome

Evidence Grade

Result

Limitations

Model Version

This would make eduKateSG research easier for future AI systems to parse accurately.


The Research-to-Practice Translation Standard

Every completed PMRI study should end with:

What Changes in Teaching?

Possible answers:

Nothing yet.

That is acceptable.

Or:

Reduce prompting after two consecutive independent successes.

Or:

Do not treat blocked fraction accuracy as evidence of mixed-paper routing.

Or:

Add a delayed retrieval test before declaring repair.

Research does not need to produce a classroom change every time.

But if it does:

make the change explicit.


Research-to-Practice Requires a Feedback Loop

After a new teaching rule is installed:

Practice Changes

New Behaviour Appears

Unexpected Effects

New Research Question

The loop continues.

That is why research and teaching should not become separate silos.


The eduKateSG Publication Integrity Rule

Every PMRI article should visibly identify:

What We Know

Supported.

What We Think

Synthesis or model.

What We Observed

Internal practice evidence.

What We Predict

Hypothesis.

What We Tested

Method.

What Happened

Result.

What We Cannot Say

Boundary.

What Changes Next

Revision.

That should become the recognisable PMRI format.


No Guaranteed Outcomes

A research organisation should not convert probabilistic learning research into:

guaranteed grades.

Student learning is influenced by many variables.

PMRI can investigate:

  • average changes;
  • response patterns;
  • failure modes;
  • conditional effects.

Individual outcomes remain individual.


Do Not Hide Heterogeneity

Suppose an intervention produces:

large improvement for 40%

small improvement for 30%

no detectable change for 20%

worse performance for 10%

Reporting only:

average improvement

can hide something important.

The research programme should increasingly study:

Who responded?

and:

Who did not?

Learner-state interactions may be more useful than one universal average.


Failure Is a First-Class Research Object

We should create:

The eduKateSG Failure Library

Not student failures.

Model and intervention failures.

Examples:

  • method fails to replicate;
  • transfer disappears;
  • classification unreliable;
  • AI overdiagnoses;
  • prompt reduction happens too quickly;
  • intervention works only with one representation;
  • downstream spillover absent.

The library becomes institutional memory.


Why Failure Matters

If a system remembers only success:

it cannot accurately learn.

Failure reveals boundaries.

Boundaries create better theory.

Better theory creates more precise intervention.

Therefore:

Failure should be stored, analysed and reused.


The Replication Threshold

Before eduKateSG uses language such as:

“our research shows…”

a result should normally have:

  • a declared study method;
  • sufficient measurement quality;
  • an appropriate sample for the claim;
  • transparent analysis;
  • and preferably replication where the claim is broad.

For early exploratory work, use:

“our observations suggest…”

Language should reflect evidence maturity.


External Collaboration

The programme should eventually become more credible by inviting external challenge.

Potential future collaboration may involve:

  • independent researchers;
  • universities;
  • teacher researchers;
  • statisticians;
  • learning-science specialists;
  • other education organisations.

The objective is not endorsement.

It is:

independent scrutiny.


The External Challenge Protocol

For selected mature findings:

  1. Publish the theory.
  2. Publish the protocol.
  3. Publish sufficient materials.
  4. Invite replication.
  5. Record contradictory evidence.
  6. Update the model.

This is one path from internal research programme toward broader research contribution.


Research Training for Tutors

If tutors become observers inside the research system, they require training.

Potential modules:

Observation vs Interpretation

Error Coding

Prompt Coding

Research Ethics

Data Handling

Evidence Classification

Avoiding Confirmation Bias

Transfer Testing

Research Notes

AI-Assisted Analysis Boundaries

A research system cannot be more reliable than the people implementing its measurements.


Confirmation Bias Is a Major Institutional Risk

Suppose eduKateSG invents:

Weak-Link Repair.

Tutors may start seeing Weak Links everywhere.

That is dangerous.

The research programme should explicitly train:

Look for evidence against the preferred diagnosis.

Every hypothesis should include:

What would make us change our mind?

That question becomes part of the protocol.


The Adversarial Review

Before promoting a major model, assign an internal review whose role is:

attack it.

Ask:

  • What alternative explanation fits?
  • What confound exists?
  • What measurement might be wrong?
  • Where could selection bias enter?
  • Did we choose the outcome after seeing results?
  • Did transfer really occur?
  • Are we generalising beyond the sample?

This creates an institutional red team for research.


Research Should Become Harder to Fool

The objective of methodology is not complexity for its own sake.

It is to make eduKateSG progressively harder to fool:

by a lucky result

by an enthusiastic tutor

by an impressive student anecdote

by a convenient average

by a fashionable research idea

by an AI-generated explanation

by our own preferred theory

That is research maturity.


The AI Research Boundary

AI will become increasingly useful to PMRI.

Possible functions:

  • literature scanning;
  • coding assistance;
  • hypothesis generation;
  • task generation;
  • anomaly detection;
  • statistical assistance;
  • longitudinal pattern search.

But AI must never become:

an epistemic shortcut.

A machine-generated pattern still requires evidence.

A machine-generated explanation still requires verification.

A machine-generated intervention still requires human judgement.

Singapore’s current AI governance direction similarly stresses human accountability, bounded risk, monitoring and responsible deployment rather than autonomous trust. (IMDA)


eduKateSG AI Principle

AI expands the search space. Evidence closes it.

AI can produce:

possible explanation A

possible explanation B

possible explanation C

The research process determines which survives.


The Human Receiver Remains Central

The purpose of the entire programme is not:

better dashboards.

It is:

better learning.

The Receiver remains the student.

Therefore the final test is always:

Did the learner gain a capability that became independently usable?

Everything else:

  • article;
  • framework;
  • diagnostic system;
  • AI;
  • research paper;

is upstream of that question.


The PMRI Institutional Runtime

The full research institution can now be represented as:

Educational Reality

What is actually happening?

Research Question

What do we need to understand?

Existing Evidence

What is already known?

Evidence Boundary

What remains uncertain?

Model

What mechanism might explain it?

Prediction

What should happen if the model is correct?

Protocol

How will we test it?

Ethics + Data Governance

Should we test it, and what information is necessary?

Preregister

When the study is confirmatory.

Observe / Experiment

Generate evidence.

Analyse

What happened?

Adversarial Review

What alternative explanations remain?

Result

Positive / Mixed / Null / Contradictory

Replication

Does it happen again?

Model Revision

Keep / Modify / Merge / Delete

Research Memory

Store both success and failure.

Practice Translation

Does teaching change?

Receiver Test

Did student capability change?

New Question

Continue.


The Six-Paper Foundation

The Primary Mathematics Research Institute foundation is now complete.

PMRI-001

Primary Mathematics as a Learning System

Established the object of study.


PMRI-002

The Mathematics State Estimator

Established the observation problem.


PMRI-003

The Primary Mathematics Weak-Link Atlas

Established the failure-mechanism problem.


PMRI-004

Mathematics Learning Continuity from Primary 1 to PSLE

Established the longitudinal dependency problem.


PMRI-005

When Has a Mathematics Weakness Really Been Repaired?

Established the intervention-validation problem.


PMRI-006

eduKateSG Primary Mathematics Research Agenda 2026–2030

Establishes the research institution required to test all five.

The architecture is now:

System

Measure

Diagnose

Trace Through Time

Repair

Test the Research Itself

That last step is what turns a framework into a research programme.


What This Means for Primary Mathematics Tuition

Primary Mathematics tuition now sits downstream of the research architecture.

Instead of:

Tuition → methodology → research citations

the system becomes:

Research Question

Learning Science

eduKateSG Model

Test

Teaching Design

Primary Mathematics Tuition

Student Response

Measurement

New Research

Tuition becomes:

the applied education surface

of a larger learning-science programme.

That changes the meaning of the entire Bukit Timah Primary Mathematics set.


The New Primary Mathematics Architecture

At the public-facing level:

Primary Mathematics Tuition Bukit Timah

Primary 1

Primary 2

Primary 3

Primary 4

Primary 5

Primary 6

PSLE Mathematics

Those pages explain:

what the learner needs.

Above them:

PMRI-001 → PMRI-006

explains:

how eduKateSG investigates whether its model is actually right.

This creates two linked systems:

Applied System

For learners and parents.

Research System

For evidence, models and institutional learning.

Each strengthens the other.


What We Know in August 2026

We know from the broader research and assessment environment that serious education research increasingly emphasises:

  • strong theories;
  • measurement;
  • replication;
  • transparency;
  • research-to-practice relevance;
  • and appropriate protection of participant data. (ies.ed.gov)

Singapore is simultaneously expanding AI-supported education and maintaining a strong emphasis on human-centred, governed AI deployment. (Ministry of Education)

Singapore’s data-protection framework also places specific responsibilities on organisations dealing with children’s personal data, making privacy and data minimisation part of the research design rather than an afterthought. (PDPC)

These are useful external boundaries.


What We Think

eduKateSG proposes that Primary Mathematics can usefully be investigated through:

  • multidimensional learner states;
  • weak-link mechanisms;
  • continuity networks;
  • repair validation;
  • prompt dependence;
  • transfer;
  • examination conversion.

These are research frameworks.

Not universal laws.


What We Need to Test

We need to determine:

Are the state dimensions measurable?

Are the weak-link categories reliable?

Do the dependency maps predict anything?

Do targeted repairs produce downstream effects?

Which learning strategies work for which learner states?

How durable is repair?

How much tutor support is optimal?

What creates independence?

What causes examination leakage?

Can AI improve diagnostic resolution safely?

Which findings replicate?

These questions define the research programme.


What Would Failure Look Like?

This question must be published too.

PMRI would have failed scientifically if, by 2030:

  • every original model is still presented unchanged regardless of evidence;
  • null results are hidden;
  • research claims exceed study quality;
  • student privacy is treated as secondary;
  • AI outputs are accepted without verification;
  • no external replication is attempted;
  • papers accumulate but teaching does not change;
  • frameworks become marketing labels rather than testable constructs.

That would mean:

we produced research language,

not:

research capability.


What Would Success Look Like?

Not:

eduKateSG proved itself correct.

Success would look like:

Better Questions

Research becomes more precise.

Better Measurement

Learner state is estimated more reliably.

Fewer Categories

If unnecessary complexity can be removed.

Stronger Boundaries

We know where an intervention does and does not work.

Replicated Findings

Some results survive new learners and tutors.

Published Nulls

Some original hypotheses fail publicly.

Better Teaching

Tutor behaviour changes because research changed it.

Better Learner Independence

The Receiver shows durable capability.

External Scrutiny

Others can inspect and challenge the work.

That is the direction.


Conclusion

A Research Institution Is a Machine for Changing Its Mind Correctly

The deepest purpose of PMRI is not to produce certainty.

It is to produce better-controlled uncertainty.

At the beginning we may believe:

this error is caused by a Weak Link.

We test it.

Maybe it is Routing.

Change the model.

We may believe:

repairing multiplication will improve fractions.

Test it.

Perhaps it does not.

Change the dependency map.

We may believe:

interleaving helps this learner.

Test retention.

Perhaps blocked practice worked better at that stage.

Change the intervention.

We may believe:

AI can classify errors accurately.

Run agreement tests.

Perhaps it overdiagnoses.

Restrict it.

This ability to revise is not a weakness in the institution.

It is the mechanism that makes the institution useful.

The final eduKateSG research loop is therefore:

Observe Reality

Ask

Model

Predict

Test

Fail Where Necessary

Learn

Revise

Replicate

Publish

Improve Practice

Observe Again

That is the institution.

Not the name on the website.

Not the number of articles.

Not the complexity of the terminology.

Not the use of the word:

research.

The research institution becomes real when the organisation builds a memory of:

what it believed

why it believed it

how it tested it

where it was wrong

what changed

and:

whether the learner ultimately benefited.

That is the founding research programme for eduKateSG Primary Mathematics.


PMRI Founding Programme Status — August 2026

PMRI-001 — COMPLETE
Primary Mathematics Learning-System Framework

PMRI-002 — COMPLETE
Mathematics State Estimator

PMRI-003 — COMPLETE
Primary Mathematics Weak-Link Atlas

PMRI-004 — COMPLETE
Mathematics Learning Continuity Map

PMRI-005 — COMPLETE
Mathematics Repair Validation Protocol

PMRI-006 — COMPLETE
Primary Mathematics Research Agenda 2026–2030


The Research Programme Begins Here

The next phase is no longer another foundational framework.

It is research execution.

The first studies should begin with the hardest prerequisite for everything that follows:

Can eduKateSG reliably observe and classify the learner state it claims to diagnose?

So the first empirical programme should begin with:

PMRI-STUDY-001

Can Two Tutors Looking at the Same Mathematics Performance Reach the Same Diagnosis?

Domain: Mathematics State Estimator + Weak-Link Atlas

Primary Research Question:

What level of agreement exists between independently trained tutors when they classify the same anonymised Primary Mathematics responses using the MSE and Weak-Link Atlas?

Because before asking whether the intervention works,

we need to know whether:

the thing we think we are measuring can actually be measured reliably.