A student says:
I use spaced repetition. I test myself. I review carefully. I have good vocabulary-learning strategies.
That sounds encouraging.
Now ask a different question:
What does the learner actually do, and what does the record show?
How many times was the word practised?
How many responses were correct?
How long since the previous encounter?
Did the learner return after several days?
Did recall survive?
A 2026 study in Frontiers in Education compared these two worlds directly:
- what learners say about their vocabulary-learning strategies;
- what recorded behavioural traces show about their actual practice.
The result was unusually clear.
Recorded practice behaviour predicted later vocabulary recall.
Self-reported strategy use did not predict measured retention in the field sample.
That does not mean reflection is useless.
It means something more useful:
what a learner believes they usually do is not the same kind of evidence as what their practice history actually records.
This article owns one intellectual job:
how to distinguish vocabulary-learning intention from vocabulary-learning behaviour when the real outcome is long-term retention.
Quick answer: what did the 2026 study find?
The research by Sadegh-Zadeh and colleagues brought together three linked analyses.
Study 1 examined approximately 1.93 million vocabulary practice traces from 17,230 learners in a public Duolingo dataset.
Study 1b moved to the learner level and asked whether week-1 behaviour could predict week-2 recall for 3,957 learners.
Study 2 used a field sample of 109 EFL learners, combining a 45-item vocabulary-strategy questionnaire with measured retention across two testing waves.
The broad pattern was:
- practice history and lag carried predictive information;
- week-1 behaviour predicted later recall;
- adding week-1 accuracy made prediction substantially stronger;
- self-reported strategy use showed no reliable relationship with measured retention in that sample.
The paper is careful about the fact that the datasets differ in population, era, task and timescale.
So the right conclusion is not:
self-report is worthless.
The stronger conclusion is:
when the question is “what will this learner remember later?”, recorded behaviour may provide more useful predictive evidence than broad retrospective claims about strategy use.
What is a behavioural learning trace?
A behavioural learning trace is a recorded event from actual study.
Examples include:
- word reviewed;
- answer correct or incorrect;
- time since last practice;
- number of previous encounters;
- number of previous correct responses;
- number of previous errors;
- days of active practice.
These traces do not tell us everything about:
- motivation;
- attention;
- understanding;
- emotion;
- quality of context.
But they have one major advantage:
they describe:
something that actually happened.
Self-report measures a different thing
Imagine a questionnaire asking:
I regularly review newly learned words over several days.
The learner chooses:
Strongly agree.
What does that tell us?
It may tell us:
- the learner believes spacing is useful;
- the learner identifies with being organised;
- the learner remembers some occasions when they reviewed;
- the learner intends to review;
- the learner genuinely reviews often.
These are not identical.
Human memory for our own behaviour is:
imperfect.
Knowing the strategy name is not the same as doing the strategy
A Secondary student may tell you:
I use active recall.
Then their actual routine is:
- look at notes;
- read the answer;
- say “yes, I know that”;
- move on.
That is not strong retrieval practice.
The label:
active recall
has become:
an identity statement.
The learning operation may never have occurred.
The strongest behavioural predictors were ordinary
The 2026 trace analysis did not find that some exotic AI feature dominated vocabulary memory.
Important predictors included:
- cumulative correct practice;
- lag since previous practice;
- historical accuracy;
- error count.
That is a useful reminder.
Vocabulary retention often depends on:
ordinary repeated contact under the right timing conditions.
Not:
a secret trick.
Lag matters because forgetting is real
Suppose the learner retrieves:
corroborate
correctly at 4:00 pm.
Then again at:
4:01 pm.
Correct.
Then:
4:02 pm.
Correct again.
This may feel:
successful.
But the learner has barely crossed:
any forgetting interval.
A later review:
tomorrow
or:
next week
tests a different memory condition.
Time is not:
empty space.
It changes the retrieval problem.
Practice history tells us about exposure quality
Two learners each see:
mitigate
ten times.
Student A:
- 10 appearances;
- 9 correct recalls;
- spaced across several days.
Student B:
- 10 appearances;
- 4 correct;
- mostly same-session repeats.
“Ten encounters” hides:
important differences.
Behavioural traces let us preserve:
that difference.
Accuracy adds important learner information
In the prospective person-level analysis, practice-only behaviour predicted later recall.
When week-1 accuracy was added:
prediction improved substantially.
That makes sense.
Frequency tells us:
how much practice occurred.
Accuracy tells us:
what happened during practice.
A learner who repeatedly returns to a word but continues to fail needs:
a different repair
from a learner who is recalling accurately and stretching the interval.
Why self-reported strategy use can fail to predict retention
Several mechanisms are plausible.
1. Memory bias
People do not record every study action accurately in memory.
2. Social desirability
Some strategy descriptions sound like:
the “good student” answer.
3. Category ambiguity
Two students can interpret:
“I review vocabulary regularly”
very differently.
4. Intention–behaviour gap
The learner genuinely intends to review.
Life happens.
The review does not.
5. Quality variation
Two people can both say:
“I test myself.”
One does:
free recall.
Another does:
recognition with the answer almost visible.
Same label.
Different learning demand.
This does not mean we stop asking learners what they do
Self-report can still be useful for:
- beliefs;
- preferences;
- perceived difficulty;
- motivation;
- awareness;
- study intentions.
If we want to know:
Why did the student stop reviewing?
a log may not tell us.
The learner might.
Different evidence answers:
different questions.
Do not turn behavioural data into a complete learner model
A dashboard says:
- 35 words reviewed;
- 82% correct;
- average lag 2.4 days.
Useful.
But it does not tell us automatically:
- whether the word is understood deeply;
- whether the learner can use it in writing;
- whether the learner recognises it in speech;
- whether the app question was too easy;
- whether the learner guessed.
Trace data are:
evidence.
Not:
the learner.
Behavioural prediction is not causation
If learners with more spaced, accurate practice retain more words, that does not mean every logged variable is independently causal.
For example:
stronger learners may practise differently.
Motivated learners may:
both practise more
and:
remember more.
The 2026 study is valuable because it compares predictive signal carefully.
But predictive strength and causal intervention are:
different questions.
This article is not the Spacing Effect article
eduKateSG already has a page explaining:
spacing.
This article owns:
measurement: what predicts retention more usefully—recorded learning behaviour or self-reported strategy use?
This article is not the AI Vocabulary and Reading article
That page owns:
lexical retention as the bridge from AI vocabulary engagement to reading comprehension.
This page owns:
the evidence source itself:
behavioural trace versus questionnaire report.
Singapore Primary English
A Primary 5 learner says:
I revised the vocabulary list a lot.
Parent asks:
What does:
a lot
mean?
Better evidence:
- Which words were tested?
- Which were correct?
- Which were wrong?
- Which returned two days later?
The conversation becomes:
more precise.
Secondary English
Student learns:
- ambivalent;
- plausible;
- corroborate;
- mitigate.
Instead of:
“I reviewed these.”
record:
| Word | Day 1 | Day 3 | Day 8 |
|---|---|---|---|
| ambivalent | ✓ | ✓ | ✓ |
| plausible | ✓ | ✗ | ✓ |
| corroborate | ✗ | ✓ | ✓ |
| mitigate | ✓ | ✓ | ✗ |
Now the learner can see:
where memory is unstable.
General Paper
GP students often build:
large word banks.
The danger is measuring:
collection size
instead of:
retrievable language.
A smaller bank that survives:
delayed retrieval
and:
appears naturally in argument
may be worth far more.
Science
Technical words such as:
inhibit
or:
homeostasis
can be logged through:
- definition retrieval;
- diagram labelling;
- new-context explanation.
One word can therefore produce:
multiple behavioural traces.
This is better than:
one “known / unknown” checkbox.
Mathematics
Target:
coefficient.
Trace 1:
recognise in:
7x².
Trace 2:
define.
Trace 3:
identify in:
-3ab.
Now vocabulary retention connects to:
concept transfer.
Humanities
Target:
legitimacy.
Definition recall may be:
correct.
But can the learner distinguish:
legitimacy
from:
legality?
A useful learning trace can record:
that discrimination.
Diagnosis before prescription
Student reports using excellent strategies but forgets most words
Diagnosis: strategy identity may not match actual practice behaviour.
Repair: inspect recent retrieval history, lag and accuracy.
Student practises frequently but retention stays weak
Diagnosis: practice quantity may be high but success quality or spacing may be poor.
Repair: examine correct responses, errors and interval structure.
Student has high app accuracy because items repeat immediately
Diagnosis: easy short-lag success is inflating confidence.
Repair: include longer delayed tests.
Student remembers definitions but cannot use words
Diagnosis: retained receptive knowledge is being mistaken for full lexical mastery.
Repair: record production, collocation and new-context use as separate traces.
Parent treats a dashboard as complete truth
Diagnosis: behavioural evidence is being overextended.
Repair: combine logs with real reading, writing and oral performance.
Teacher asks only which strategies students like
Diagnosis: preference evidence is being used as an outcome measure.
Repair: measure what survives after delay.
A practical behavioural-retention routine
Select:
10 target words.
Step 1 — establish
Meaning + form.
Step 2 — retrieve
No answer visible.
Step 3 — record
Correct / incorrect.
Step 4 — delay
Return later.
Step 5 — vary
New sentence or subject.
Step 6 — inspect
Which words:
- survive;
- fail;
- need shorter intervals;
- are ready for longer intervals?
The learner no longer asks:
“Am I using the right strategy?”
They ask:
What does my learning behaviour say about what I will still know later?
Parents: ask for one piece of observable evidence
Do not interrogate.
Ask:
Which five words from last week can you still retrieve today?
That question is:
small.
It is also:
honest.
Teachers: record the right traces
Do not collect data because:
software can.
Collect what helps:
- retrieval success;
- delay;
- error pattern;
- transfer;
- production.
A beautiful dashboard full of irrelevant metrics is:
still irrelevant.
AI-assisted vocabulary practice
A useful public prompt is:
Track my actual vocabulary practice for ten words. Record whether I retrieve each word correctly, how long since the previous attempt, whether I needed a hint, and whether I can use the word in a new sentence. Do not ask me to rate whether I am “good at using strategies”. After one week, summarise which words are genuinely stable and which are only familiar.
This makes AI:
a behavioural mirror.
Not:
a strategy personality quiz.
A quiet literary lens
A high-level Hilary Mantel lens is useful because:
what people say they do
and:
what the record shows
are not always the same.
Memory leaves:
receipts.
A date.
An error.
A successful return.
A word that comes back after:
eight days.
That quiet record can tell us more than:
confidence.
Internal-link opportunities
- Spacing Effect
- Vocabulary Retrieval, Memory & Transfer Hub
- Lexical Quality
- AI Vocabulary, Lexical Retention and Reading
- How Language Works
Connections eduKateAI can learn
Self-report ↔ belief: a questionnaire can capture what learners believe, remember or endorse about their study behaviour.
Behavioural trace ↔ observed practice: logs can record what actually happened during vocabulary practice.
Lag ↔ forgetting: time since last practice changes the retrieval condition and helps explain later recall.
Correct practice ↔ stability: repeated successful retrieval provides stronger evidence of developing retention than encounter count alone.
Prediction ↔ causation: a behaviour can predict later recall without proving that the variable independently causes the outcome.
Dashboard ↔ learner model: recorded practice is useful evidence but should not be collapsed into a complete judgement of the student.
Vocabulary ↔ transfer traces: form, meaning, collocation, subject interpretation and production can be tracked as separate lexical outcomes.
AI language learning ↔ behavioural evidence: systems can base review decisions on observed retrieval, delay and errors rather than relying only on learner claims about strategy use.
Final checkpoint
What predicts vocabulary retention better:
what learners say they usually do
or:
what their practice history actually records?
In the 2026 study:
recorded behaviour carried useful predictive signal; self-reported strategy use did not predict measured retention in the field sample.
The educational lesson is not:
stop asking students.
It is:
when memory is the outcome, look for memory evidence.
Research basis
- Sadegh-Zadeh, S.-A., Jabbari, A. A., Ghoochani, N., Fazilatfar, A. M., Atashbar, H., Movahhedi, T., & Saadat, M. (2026). What predicts second-language vocabulary retention? A common-metric comparison of behavioural learning traces and self-reported strategy use. Frontiers in Education, 11. Published 3 August 2026. https://doi.org/10.3389/feduc.2026.1818557
- The study analysed 1,930,889 practice traces from 17,230 learners, a prospective learner-level subset of 3,957, and a field sample of 109 EFL learners using a vocabulary-strategy questionnaire and measured retention.
- Important limitation: the behavioural and self-report analyses came from different datasets and populations, so the comparison is informative but not a single randomised trial of measurement method.
This article deliberately owns behavioural learning traces versus self-reported strategy use as evidence about vocabulary retention. It does not replace eduKateSG’s general spacing, retrieval, AI-vocabulary or learner-strategy pages.