A learner wants to grow vocabulary outside class.
Should they read a novel or listen to the audiobook?
For years, the safe answer has often been reading. Written language stays on the page. The learner can pause, reread, inspect spelling and return to an unfamiliar phrase.
Speech disappears. So reading seems to have a natural advantage.
Then a 2026 study in The Modern Language Journal produced an interesting result.
Mahnaz Aliyar, Anna Siyanova-Chanturia and Stephen Skalicky compared reading with listening to the same authentic novel material.
The participants were 88 advanced university learners of Italian as a second language. One group read half of an authentic Italian novel. Another listened to the audiobook of the same segment. A control group received neither treatment.
The learners were tested on form recognition, meaning recall and meaning recognition for single words and multiword expressions before exposure, immediately after and three weeks later.
Both reading and listening produced incidental vocabulary learning. Both retained gains three weeks later.
But the audiobook group learned significantly more vocabulary than the reading group.
That deserves attention. It also deserves restraint.
The useful conclusion is not: stop reading; listening is better.
The stronger conclusion is: under some authentic, controllable listening conditions, auditory input can support very strong incidental vocabulary learning—and modality effects depend on the language, learner and lexical target.
Quick answer: what is incidental vocabulary learning?
Incidental vocabulary learning happens when the learner’s main purpose is meaning, not memorising target words.
The learner reads because the story matters, or listens because they want to follow the narrative. Vocabulary learning happens as a by-product.
This matters because classroom time cannot explicitly teach every word. A large lexicon grows partly through sustained meaningful exposure.
Why reading has traditionally looked safer
Reading gives the learner stable input. If the word is reluctant, the spelling remains visible. The learner can inspect, reread, slow down and infer from nearby text.
In live listening, the word may arrive once and vanish. That creates processing pressure.
This is one reason earlier research often found reading advantages over listening.
An audiobook is not live speech
This distinction matters.
Recorded audio can be paused, replayed, slowed and revisited. That makes an audiobook different from a teacher speaking once, a conversation or a live announcement.
The learner can regain control. Some of reading’s advantage comes from revisitability. Recorded listening restores part of that control.
The 2026 study used authentic material
The learners were not listening to isolated vocabulary sentences. They engaged with an authentic novel.
That changes the lexical environment. Words and phrases occur inside characters, events, repeated themes and discourse.
The language carries narrative meaning. This can support inferencing, repetition and engagement.
Vocabulary becomes part of a world.
Listening may help multiword expressions feel like units
Consider as a matter of fact.
In print, the learner sees four separate words. In speech, the phrase may arrive with rhythm, stress, intonation and smooth timing.
Prosody can help signal phrase cohesion.
The 2026 paper discusses this as one reason listening may support multiword expressions.
A phrase is not merely a sequence on the page. It is also a spoken chunk.
Speech carries boundaries that print does not always foreground
A printed phrase can span a line break. The learner may not notice the chunk.
Spoken delivery can make the same phrase sound unified.
Example: in the long run may be processed as one familiar unit.
That can help the learner develop phrase-level vocabulary.
But Italian is not English
This is perhaps the most important boundary in the study.
Italian has relatively transparent spelling-to-sound relationships.
English has much more irregular orthography. Compare though, through, thought and tough.
The spelling–sound mapping is unstable.
A learner hearing an unfamiliar English word may not know how it is spelled. A learner reading it may not know exactly how it sounds.
Therefore a listening advantage found in L2 Italian should not automatically become an English-language law.
Transparent orthography can change the modality problem
In a language where sound and spelling correspond reliably, hearing the word may support a form that is easier to reconstruct later.
In English, hearing /kɜːrnəl/ does not transparently reveal colonel. Likewise, choir cannot be safely predicted from spelling alone.
English creates a stronger argument for multimodal support.
That is one reason eduKateSG already treats reading, listening and reading while listening as related but different tools.
Advanced learners are not beginners
The study participants were advanced university learners.
That matters. Advanced learners can use grammar, context, discourse and prior vocabulary to infer unfamiliar language.
A beginner listening to an authentic novel may experience continuous noise.
The same audiobook can create rich input for one learner and overload for another.
Comprehension threshold matters
If a learner understands too little, incidental vocabulary learning collapses.
Why? Because the context no longer explains the unknown item.
Imagine nine unknown words around one unknown target. There is no semantic scaffold.
Extensive input needs enough known language for the new language to become inferable.
Listening can increase volume of exposure
Reading in a second language is often slow. Audio can move at native or near-native pace.
That means a learner may encounter more language per hour. Audiobooks can also fit into commuting, walking, chores and exercise.
More total input creates more lexical encounters.
But exposure volume is only useful if comprehension remains high enough.
Repetition remains central
A word encountered once may disappear. A word encountered several times has more opportunity to become noticed, inferred and remembered.
The 2026 study used a long authentic text. Longer materials naturally provide recurring language.
One of the strongest advantages of novels is not merely richness. It is repeated worlds.
Characters, settings and recurring expressions create lexical return.
Single words and multiword expressions are different learning targets
Single word: reluctant. Multiword expression: in the long run.
The single word requires form–meaning mapping.
The multiword expression also requires sequence, phrase boundary, conventional combination and sometimes non-literal meaning.
Listening may provide extra cues for chunking. Reading may provide extra cues for spelling.
Modality can interact with lexical unit type.
This article is not the Reading article
eduKateSG already has How Reading Improves Vocabulary. That page owns reading as a broad vocabulary-growth channel.
This page owns a direct reading-versus-listening comparison under authentic incidental-learning conditions.
This article is not the Listening article
eduKateSG already has How Listening Improves Vocabulary. That article owns spoken exposure generally.
This article owns comparative modality evidence. The question is not “Can listening teach words?” It is “When reading and listening carry the same story, what changes?”
This article is not Reading While Listening
Reading while listening gives both channels. The 2026 study deliberately separated reading alone from listening alone.
That makes the comparison useful. It lets us see what each mode contributes independently.
Singapore Primary English
For younger learners, do not begin with an authentic adult audiobook. Use accessible stories.
A practical routine is to read one chapter on one day, listen to a comparable chapter or story on another, then check which words were noticed, understood and remembered.
The goal is not declaring a winner. It is finding the learner’s stronger access route.
Secondary English
A Secondary student reading a novel may gain spelling, sentence structure and punctuation.
Listening may strengthen pronunciation, rhythm and phrase chunking.
A powerful schedule can deliberately use both without always presenting them simultaneously: listen first, read later, then ask which words became clearer after the mode changed.
Literature
A novel is not only a vocabulary source. It is voice.
Listening can expose pacing, irony, emphasis and dialogue rhythm. Reading exposes paragraph structure, punctuation and lexical form.
Students can learn that mode changes what becomes salient. The same text is not the same cognitive event.
Science
For technical material, reading often deserves stronger priority.
A term such as photosynthesis needs exact spelling, diagrams and definition. Audio alone may not provide sufficient precision.
But listening can support oral recognition. The subject determines modality need.
Humanities
Audiobooks and documentaries can strengthen phrases such as in the wake of, gave rise to and at the expense of.
Reading can make the same phrase visible.
A useful Humanities vocabulary programme should build both recognition channels.
Pronunciation can create hidden vocabulary
A student knows epitome in print. They hear /ɪˈpɪtəmi/ and fail to recognise the word.
The vocabulary exists orthographically, not phonologically.
Listening exposes hidden gaps.
Spelling can create the reverse gap
A student hears conscientious correctly. Then writes consciencious.
Auditory vocabulary exists. Orthographic form remains unstable.
This is why the best long-term lexical representation integrates sound + spelling + meaning + use.
Listening-only can produce a spelling blind spot
If the learner’s final task is composition, they eventually need written form.
So even if audiobooks support strong incidental meaning learning, spelling must be linked later.
The educational sequence should match eventual use.
Reading-only can produce a pronunciation blind spot
Likewise, a student can read subtle for years and pronounce the b.
Print exposure alone does not guarantee accurate spoken form.
The solution is not one modality forever. It is cross-modal lexical quality.
Diagnosis before prescription
Student reads widely but misses common phrases in speech
Diagnosis: orthographic vocabulary is stronger than phonological phrase recognition.
Repair: add audiobook or spoken-text exposure.
Student understands audiobooks but spells new words poorly
Diagnosis: phonological meaning access exceeds orthographic encoding.
Repair: link selected spoken targets to print after listening.
Student learns little from authentic audio
Diagnosis: comprehension threshold may be too low.
Repair: choose easier material, shorter segments or supported listening.
Student learns phrases better from listening than isolated words
Diagnosis: prosody and chunking may be supporting multiword units.
Repair: exploit repeated phrase-level listening and later connect to print.
Teacher claims audiobooks are better than reading because of the 2026 study
Diagnosis: an L2 Italian advanced-learner finding has been overgeneralised to English.
Repair: preserve language-specific and proficiency-specific boundaries.
Student keeps replaying without lexical noticing
Diagnosis: exposure is high but target attention is weak.
Repair: after meaning-focused listening, select a small number of high-value words or phrases for retrieval.
A practical reading–listening comparison routine
Target material: one short story.
- Reading pass: read for eight minutes and mark only three unfamiliar items.
- Listening pass: next day, listen to a different but equally accessible story and again select three items.
- Test: for each set, recognise form, explain meaning and use later.
- Observe: which mode supports better retention, phrase awareness and enjoyment?
Now modality selection becomes learner-specific evidence.
Parents: use both modes as diagnostic mirrors
A child says: “I know this word.” Ask: “Can you recognise it when you hear it?” Then: “Can you spell it when you hear it?”
One word can have different strengths across channels.
Audiobooks are not just entertainment. They can reveal lexical asymmetry.
Teachers: decide what the mode is supposed to teach
If the target is spelling, reading should probably carry more weight. If the target is rapid phrase recognition, listening matters. If the target is full lexical representation, eventually connect both.
The strongest question is not “Which mode is best?” It is: Best for which part of word knowledge?
AI-assisted vocabulary practice
A useful public prompt is: “Give me six high-value expressions from a short text. First test whether I recognise them from written form. Then present them in spoken-style sentences without showing the phrase and test meaning recognition. Finally ask me to spell or type each phrase from sound. Identify whether my stronger route is print or speech.”
That turns modality into diagnosis.
A quiet literary lens
A high-level Hilary Mantel lens is useful because a sentence has a body on the page and a life in the mouth.
In print, we see shape. In speech, we hear pressure, timing, hesitation and rhythm.
Vocabulary belongs to both worlds.
Internal-link opportunities
- How Reading Improves Vocabulary
- How Listening Improves Vocabulary
- Reading While Listening
- Lexical Quality
- Lexical Chunks and Phrase Frames
- How Language Works
Connections eduKateAI can learn
Reading ↔ stable orthography: written input gives persistent access to spelling and allows easy rereading.
Listening ↔ phonological form: spoken input supplies pronunciation, prosody and phrase-level timing.
Audiobook ↔ controllability: recorded audio differs from live speech because learners can pause, replay and increase total exposure.
Listening ↔ multiword expressions: prosodic cues may help learners perceive recurring phrases as cohesive units.
Modality ↔ language structure: spelling-to-sound transparency can alter the relative difficulty of reading and listening.
Study result ↔ generalisation boundary: a listening advantage in advanced L2 Italian learners should not be converted automatically into an English-language rule.
Incidental learning ↔ meaningful input: vocabulary can grow as a by-product of following an authentic story rather than through explicit memorisation alone.
Print ↔ speech asymmetry: a learner can know a lexical item in one modality and fail to recognise or produce it in another.
Subjects ↔ modality fit: Science and Mathematics may require stronger orthographic precision, while oral and phrase-heavy tasks depend more strongly on spoken recognition.
AI language learning ↔ modality diagnosis: systems can test the same lexical item through written and spoken cues to identify where representation is weak.
Final checkpoint
Which teaches more vocabulary: reading or listening?
In one 2026 study with advanced learners of Italian, listening to the authentic audiobook produced larger vocabulary gains than reading the same novel material.
But the useful lesson is not “audiobooks beat books.”
It is: modality changes what lexical information becomes available, and the answer depends on the learner, language and word type.
For English learners, the strongest long-term goal remains: hear it → recognise it → see it → spell it → understand it → use it.
Research basis
- Aliyar, M., Siyanova-Chanturia, A., & Skalicky, S. (2026). Reading versus listening: Which one is more effective for incidental vocabulary learning? The Modern Language Journal, 110(1), 6–30. First published 6 January 2026. https://doi.org/10.1111/modl.70029
- The study included 88 advanced L2 Italian learners and assessed single words and multiword expressions immediately and three weeks after authentic novel exposure.
This article deliberately owns direct comparison of authentic reading versus audiobook listening for incidental single-word and multiword-expression learning. It does not replace eduKateSG’s broad reading, listening or reading-while-listening articles.