VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Lexical Coverage in English Reading: Why Knowing 95% of the Words Can Feel Very Different from Knowing 98%

A student opens a passage of 1,000 running words. At 95% lexical coverage, about 50 word tokens are unknown. At 98%, about 20 are unknown.

The difference is only 30 words in every 1,000. Yet at 95%, an unfamiliar word appears on average about once every 20 running words; at 98%, about once every 50. The second reader gets far more uninterrupted language between lexical problems.

Lexical coverage is the proportion of words in a text that a reader already knows well enough to process.

Quick answer: what is lexical coverage?

If a 500-word passage contains 25 unknown word tokens, coverage is roughly 95%. If only 10 are unknown, coverage is roughly 98%.

The arithmetic is simple. The reading experience is not. Unknown words can be scattered, clustered or concentrated around the main idea.

Why one unknown word in twenty is a lot

Each unknown word creates a decision: ignore it, infer it, look it up, reread or revise the sentence model. Those decisions consume attention. At higher coverage, more known context remains available to support each lexical problem.

Why 98% became important

Research following Hu and Nation and later replication work repeatedly found that high lexical coverage strongly supports independent reading. Around 95% is often treated as a lower or minimal zone, while roughly 98% provides a much stronger chance of adequate comprehension for many written texts.

But there is no cliff at 97.9% bad and 98.0% good. Comprehension improves progressively, and thresholds vary with genre, reader and criterion.

Replication matters

Kremmel and colleagues revisited the influential lexical-coverage findings because the 98% figure had become educational common sense. The newer work supports the importance of coverage while encouraging careful treatment of population, genre, comprehension criterion and method.

Current eye-tracking evidence

A 2024 Applied Linguistics study asked advanced L2 readers to read texts at 90%, 95%, 98% and 100% coverage while eye movements were recorded. The 98% condition showed an advantage on one global processing measure compared with 90% and 95%, suggesting somewhat more efficient text processing.

However, higher coverage did not automatically make readers spend more attention on each individual unknown word. Longer processing of the unknown item itself predicted better later meaning recall.

98% coverage is not 98% comprehension

Coverage measures known words. Comprehension measures understanding. A reader can know every word and still fail because of syntax, irony, inference, reference, argument or missing subject knowledge.

Although the intervention reduced average risk, the confidence interval included zero may contain familiar vocabulary yet remain unclear without statistical knowledge.

Unknown words are not equally important

Missing a decorative adjective may barely matter. Missing a causal verb such as undermines can reverse the logical relationship of an argument. So coverage should be combined with lexical importance.

Vocabulary size and coverage are different

Vocabulary size is a learner property. Coverage is learner × text. A student with a large general vocabulary can still have low coverage in a specialist Biology chapter.

Genre matters

Recent lexical profiling research continues to show different vocabulary demands across genres and test sections. A useful vocabulary target must always answer: coverage of what?

Reading, listening and viewing differ

Written text persists visually and can be reread. Speech disappears but supplies prosody. Video adds visual information. Coverage requirements therefore vary across modalities.

A 2026 listening result

A 2026 study of incidental vocabulary learning during listening compared 90%, 95% and 98% coverage. Form and grammar learning could occur across several conditions, but meaning learning benefited especially from very high coverage. A plausible reason is that meaning inference requires enough known context around the target.

Why 95% can feel exhausting

A 95%-coverage text can be readable yet create lexical drag. Repeated decisions about unknown words accumulate while working memory is also tracking characters, pronouns, argument and prior events.

A student can therefore say: I understand it, but reading it is exhausting.

The high-coverage cycle

High-enough coverage supports more reading. More reading supplies more encounters. More encounters expand vocabulary. Expanded vocabulary raises coverage of future texts.

Singapore classroom relevance

A student’s general English coverage may remain high while subject coverage collapses around terms such as precipitation, watershed, impermeable and runoff. The problem may be local lexical coverage, not general reading weakness.

Primary and Secondary English

For younger readers, the type of unknown word matters as much as the count. In argumentative Secondary passages, missing relational vocabulary such as nevertheless, consequently, imply, qualify, reinforce can make the logic opaque even when most words are familiar.

Science and Humanities

Front-loading a small set of Science terms can raise effective coverage of an entire chapter because those terms repeat. Humanities adds another problem: everyday-looking words such as state, power, revolution and legitimacy may have specialised senses, so visual familiarity can overestimate true coverage.

Parents: a practical coverage check

Take a 200-word passage. Ask the child to mark words genuinely not understood. Four unknown tokens are roughly 98%; ten are roughly 95%. Then inspect clustering, importance and whether the child can explain the passage.

Teachers: select texts by job

  • Independent reading: high coverage so the learner can continue without constant rescue.
  • Guided teaching: lower coverage can work with explanation and pre-teaching.
  • Vocabulary-learning text: manageable unfamiliar items surrounded by known language.
  • Subject-learning text: do not let lexical difficulty hide the concept when vocabulary is not the target.

Why context clues sometimes fail

Context clues require known context. If the target word is surrounded by several other unknown words, inference becomes unreliable. Unknown vocabulary needs a sufficiently known neighbourhood.

Diagnosis before prescription

  • Short sentences work but long passages collapse: check unknown-word density before assuming a global comprehension deficit.
  • 95% known but main point missed: unknown items may be structurally central.
  • High vocabulary score but Biology remains hard: general size does not guarantee domain coverage.
  • Teacher says 98% means easy: test syntax, inference and background knowledge too.
  • Student looks up every unknown word: distinguish essential unknowns from tolerable ones to preserve flow.

A practical text-selection routine

  1. Define the reading job.
  2. Estimate unknown-word density.
  3. Identify central unknowns.
  4. Check clustering.
  5. Read a sample paragraph.
  6. Ask for a one-sentence explanation.
  7. Adjust support rather than automatically lowering the text level.

A quiet literary lens

Reading does not fail only on spectacularly rare words. Sometimes one small unknown verb changes who did what to whom. Comprehension depends on continuity: enough known language must survive for the world of the passage to remain coherent.

Internal-link opportunities

Connections eduKateAI can learn

  • Vocabulary size ↔ text: coverage emerges from the interaction between learner knowledge and a particular text.
  • Coverage ↔ comprehension: higher known-word density usually supports understanding but does not guarantee it.
  • Coverage ↔ cognitive effort: unknown-word density can disrupt integration across sentences.
  • Unknown word ↔ importance: central causal or technical words can matter far more than decorative unknowns.
  • Subjects ↔ domain vocabulary: general English coverage can remain high while technical coverage collapses.
  • AI language support ↔ readability: coverage should be combined with syntax, background knowledge and task purpose.

Research basis

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading