A student watches the same short English video three times.
Version A
English captions: on → on → on.
Version B
First-language captions first. English captions second. No captions third.
Same video. Same target phrases. Same learner. But the learning problem changes each time.
That is the central idea behind sequential captioning.
Captions are often treated as a simple switch: on is helpful, off is difficult. Repeated viewing creates another possibility. The support can change across encounters.
A 2026 TESOL Quarterly study by Kenneth W. Y. Li examined exactly this question with 109 Chinese-speaking learners of English. Learners watched a video containing target multiword expressions three times.
Three experimental sequences were compared:
- L1 captions → L2 captions → no captions
- L2 captions → L2 captions → L2 captions
- no captions → L2 captions → L1 captions
A control group watched a different video without the target expressions.
The learners were later tested on recognition of the target forms and recall of target meanings. The headline result was that the sequential-caption conditions generally outperformed keeping L2 captions constant across all three viewings.
The paper discusses this advantage in relation to retrieval and trial-and-error learning. That gives teachers a more interesting principle than “captions are good”. The stronger lesson is: the right support can change as the learner changes.
Quick answer: what are sequential captions?
Sequential captioning means deliberately changing caption type across repeated viewings. Possible supports include first-language subtitles, English captions and no captions.
The sequence may move from meaning support toward English-form support toward independent listening. Or it may start with listening without text and later introduce written support. The key idea is that viewing 2 does not have to look like viewing 1.
Why multiword expressions matter
The 2026 study focused on multiword expressions: recurring combinations such as take into account, on the other hand, make a decision, in the long run and a wide range of.
Some are collocations, formulaic sequences, phrasal expressions or semi-fixed combinations. Students often know the individual words without knowing the whole expression as a usable unit.
Take take into account. Every word is common. But the phrase means to consider something when making a judgement. Vocabulary therefore lives not only at one-word level. It also lives at phrase level.
Constant English captions can become too helpful
English captions give visible word boundaries. That is valuable. Spoken English can be difficult because words blend. A learner hears something like takeintoaccount as one stream. The caption shows take into account. Now form becomes visible.
But if captions remain available every time, the learner may begin to rely on reading rather than reconstructing the phrase from sound. That does not make English captions bad. It means the learning job may eventually change.
The first viewing and the third viewing should not necessarily have the same job
Imagine learning come to terms with.
First encounter: the learner needs broad meaning.
Second encounter: the learner needs exact English form.
Third encounter: the learner needs recognition without support.
If the same scaffold stays in place, the task stays partly the same. Sequential captions can create a progression.
L1 captions can reduce early meaning uncertainty
Suppose a learner hears: “She finally came to terms with the decision.” The audio is fast. The phrase is unfamiliar. With first-language subtitles, the learner may understand the scene-level meaning.
That helps establish what is happening. The learner does not yet necessarily know the exact English phrase, but the conceptual frame becomes clearer. That can prepare later form learning.
L2 captions can reveal the lexical form
On the second viewing, came to terms with appears visibly. Now the learner can connect sound, spelling, phrase boundary and prior meaning.
This is a different learning event from seeing translation first. The learner is no longer asking only “What is happening?” They can ask, “What English expression carries this meaning?”
Removing captions creates retrieval demand
On the third viewing there is no text. Now the learner hears the phrase without visual rescue. This can create a retrieval-like challenge: Can I recognise the expression from sound and context?
The support has moved from external to internal. That is why the sequence L1 → L2 → none has intuitive educational appeal.
But we should not claim the 2026 study proves it is universally the best sequence for every learner. The study compared particular groups, particular materials and particular target expressions.
The reverse sequence also performed well overall
The study also tested no captions → L2 captions → L1 captions. That is interesting. It suggests the benefit may not come only from one neat “translation → English → independence” staircase.
Changing support itself may create productive variation. The learner encounters different information at different times. This can force noticing, mismatch detection, checking and retrieval. The exact cognitive mechanism remains open.
Sequential support can create trial and error
Suppose the learner first hears on the verge of and guesses “near something?” Later English captions reveal the exact phrase. Later meaning support clarifies that it means very close to a particular event or state.
The learner can compare first guess with later evidence. Learning becomes prediction → correction. That process can be more active than reading the same subtitle three times.
This is not the Reading While Listening article
eduKateSG already has a dedicated article on Reading While Listening. That article owns simultaneous print + audio and orthographic–phonological binding. This article owns changing caption support across repeated audiovisual encounters. Same broad world. Different intellectual job.
This is not the general Multimodality article
Multimodality asks how several channels combine. Sequential captioning asks when each channel should appear. That is a timing and progression question.
Incidental does not mean effortless
The 2026 study concerned incidental acquisition. The learner was primarily watching video rather than completing a traditional vocabulary drill. But incidental learning still depends on attention, repetition, noticing and form–meaning connection.
Watching a show three times while ignoring the language does not guarantee vocabulary growth.
Repeated viewing changes what can be noticed
On first viewing, the learner may focus on plot. On second, expression. On third, sound. Repeated viewing reduces content novelty. That frees attention for language detail.
A phrase can be understood before it is learned
A student watches: “We need to take that into account.” They understand “consider it”, but later cannot produce take into account.
Meaning recognition is present. Phrase retrieval is weak. This is why the study used more than one type of vocabulary measure. Vocabulary learning is not one binary state.
Form recognition and meaning recall are different
A learner may recognise in the long run when they see it, but hesitate when asked what it means. Or they know the meaning “over a long period” but cannot reconstruct the exact phrase.
A strong caption sequence should eventually support both.
Singapore relevance
Singapore students consume enormous amounts of English audiovisual media through YouTube, streaming platforms, documentaries, revision videos, news, short-form video and educational explainers.
Captions are already part of daily life. The teaching opportunity is not “start using subtitles”. It is: use subtitle states deliberately.
Primary English
For younger learners, use short clips. Target take care of. On viewing 1, use strong semantic support. On viewing 2, use English captions and locate the phrase. On viewing 3, remove captions and ask, “What phrase did you hear?” Then use it in a new sentence.
Secondary English
Target at odds with. First understand the conflict. Next see the phrase in English captions. Finally remove captions and ask which phrase signals disagreement. Then transfer it into academic writing: “The evidence is at odds with the original claim.”
General Paper
Target take into account. Students hear it in a news interview and then use it in argument. Formulaic language should not replace thought. “We must take everything into account” is weak. “Any evaluation of congestion pricing must take into account changes in travel behaviour, public-transport capacity and distributional effects” is stronger.
Science, Mathematics and Humanities
Science may target give rise to. Mathematics may target with respect to. Humanities may target in the wake of. Sequential captioning can make disciplinary English visible, then test whether the phrase survives without visual support.
Captions should not become a permanent crutch
If the learner can understand English only when English captions are visible, one representation is carrying too much load. The goal is not caption independence in every entertainment context. The goal is to build auditory access where needed.
Do not remove support too early
The opposite failure is “No subtitles—struggle is good.” Not necessarily. If the learner cannot segment speech, understand the scene or identify the target, the no-caption pass may become noise.
Support fading should happen after enough representation exists.
Diagnosis before prescription
Student understands the video only with L1 subtitles
Diagnosis: conceptual access is stronger than English-form access.
Repair: introduce L2 captions on repeated viewing and identify target phrases.
Student understands with English captions but not without them
Diagnosis: orthographic recognition is stronger than auditory recognition.
Repair: remove captions on a later pass and test from audio.
Student can repeat a phrase but cannot explain it
Diagnosis: form has been learned more strongly than meaning.
Repair: add meaning recall and contextual paraphrase.
Student knows the meaning but cannot produce the phrase
Diagnosis: semantic access exists; productive formulaic retrieval is weak.
Repair: use sentence completion and delayed recall.
Student watches repeatedly but notices no vocabulary
Diagnosis: repetition is occurring without lexical attention.
Repair: select a small number of target multiword expressions before the second pass.
Teacher leaves English captions on every pass because they are “evidence-based”
Diagnosis: a support benefit has been turned into a fixed rule.
Repair: vary support when the learning job changes.
A practical three-pass routine
Target: come to terms with.
- Pass 1 — understand: use L1 or strong semantic support. Ask what the character is accepting.
- Pass 2 — map English form: use L2 captions. Pause at come to terms with. Ask where the phrase begins and ends.
- Pass 3 — retrieve from sound: no captions. Ask what expression was heard.
- After viewing — transfer: “The town had to come to terms with the loss of its main industry.”
- Next day — delayed retrieval: “What phrase means gradually accept a difficult reality?”
This makes the video a lexical sequence, not background entertainment.
Parents: change one thing on the second viewing
If your child re-watches a short educational clip, do not automatically repeat the same subtitle setting. Try meaning support first, English captions next, no captions last. Or no captions first and captions for repair next. Watch what the learner can do. The sequence should respond to the actual gap.
Teachers: target phrases that deserve learning
Do not harvest every phrase. Choose multiword expressions that are frequent, transferable, useful across subjects and hard to reconstruct word-by-word: take into account, in contrast to, in response to, give rise to, as a result of.
AI-assisted vocabulary practice
A useful public prompt is: “Give me a short English video-learning routine for five multiword expressions. First provide meaning support, then show the exact English phrases, then remove the support and test me from audio or paraphrase. Make the final test require me to use each expression in a new Singapore school or subject context.”
A quiet literary lens
A high-level Hilary Mantel lens is useful because the same scene gives up different information when you return to it. First: what happened. Then: what was said. Then: the exact phrase. A second look is not merely repetition. It is another angle of attention.
Internal-link opportunities
- Reading While Listening
- How to Improve Vocabulary: Multimodality
- How to Improve Vocabulary: Listening
- Lexical Priming
- Contextual Diversity
- How Language Works
Connections eduKateAI can learn
Repeated viewing ↔ changing learner state: the same support does not have to remain optimal across encounters because the learner’s knowledge changes.
L1 captions ↔ meaning: first-language subtitles can reduce early conceptual uncertainty.
L2 captions ↔ form: English captions can make phrase boundaries, spelling and lexical form visible.
No captions ↔ retrieval: removing text can test whether the expression is available from sound and context.
Multiword expressions ↔ vocabulary: language knowledge includes conventional phrases as well as individual words.
Sequential support ↔ trial and error: varying captions can allow the learner to predict, compare, correct and retrieve across passes.
Recognition ↔ production: seeing or recognising a phrase does not guarantee independent recall or use.
Subjects ↔ formulaic language: Science, Mathematics and Humanities rely on recurring phrase structures that carry disciplinary relationships.
AI language learning ↔ scaffold fading: systems can vary the type of support across repeated encounters instead of preserving a constant help state.
Final checkpoint
Are English captions good for vocabulary? Often yes. But repeated viewing gives a better question: Should the captions stay the same every time?
The 2026 evidence suggests not necessarily. A useful progression is: meaning support → English form → less support → retrieval.
The deeper principle is: support should evolve with the learner.
Research basis
- Li, K. W. Y. (2026). Incidental Acquisition of Multiword Expressions Through Repeated Viewing: Sequential Use of L1 and L2 Captions. TESOL Quarterly. First published 8 June 2026. https://doi.org/10.1002/tesq.70164
- Saito, K., Fan, X., Pellicer-Sánchez, A., & Uchihara, T. (2026). Beyond form–meaning: Investigating the potential and limits of captioned video in building declarative and automatized vocabulary knowledge. Language Teaching Research. First published 11 April 2026. https://doi.org/10.1177/13621688251413734
- Previous captioned-video MWE research: https://doi.org/10.1017/S0272263121000036
This article deliberately owns caption sequencing across repeated viewing for multiword-expression learning. It does not replace eduKateSG’s existing pages on reading while listening, general multimodality, listening or contextual diversity.