What is vocabulary lexical diversity? Lexical diversity describes how varied the vocabulary is in a piece of language: how many different words, word forms or lexical items appear relative to the total amount of language produced. It is closely related to search terms such as vocabulary richness, vocabulary variety, lexical variation, unique words and word diversity. In writing and speaking research, lexical diversity is often estimated with measures such as type-token ratio (TTR), moving-average type-token ratio (MATTR), Measure of Textual Lexical Diversity (MTLD) and other indices designed to compare how much vocabulary variation appears in a language sample.
Lexical diversity is not the same thing as vocabulary size, vocabulary depth, lexical sophistication or lexical density. A learner can know many words but repeat a small set in one essay. A text can use many different words without using rare or advanced words. A highly technical paragraph can repeat the same specialist terms and therefore show lower surface variety while still being precise and expert. These distinctions matter because a high lexical-diversity score is not automatically “better writing,” and a low score is not automatically weak vocabulary.
This guide explains lexical diversity meaning, lexical diversity examples, type-token ratio, TTR formula, MATTR, MTLD, vocabulary richness in writing and speaking, lexical sophistication, lexical density, vocabulary repetition, text length effects, student assessment and ways to improve vocabulary variety without forcing synonyms. It also connects lexical diversity back to the wider eduKate vocabulary system, where the larger goal is usable word knowledge: meaning, retrieval, precision, collocation, register and transfer—not numerical variety for its own sake.
Lexical diversity asks a narrow but useful question: how much vocabulary variety appears in this particular sample of language?
This article is a child of the canonical What Is Vocabulary? owner. The parent defines vocabulary as a complete knowledge-and-use system. This page owns one specialist measurement problem: how varied the words in a text or stretch of speech are, how that variation can be measured, and how easily the resulting numbers can be misunderstood.
Contents
- 1. The short answer
- 2. Types and tokens
- 3. Type-token ratio
- 4. Why text length changes TTR
- 5. MATTR
- 6. MTLD
- 7. Other lexical-diversity measures
- 8. Diversity vs richness
- 9. Diversity vs vocabulary size
- 10. Diversity vs vocabulary depth
- 11. Diversity vs lexical sophistication
- 12. Diversity vs lexical density
- 13. Repetition can be good
- 14. Lexical diversity in writing
- 15. Lexical diversity in speaking
- 16. Genre and topic effects
- 17. What teachers can infer
- 18. Assessment cautions
- 19. How to improve lexical diversity
- 20. Worked text samples
- 21. Diagnostic patterns
- 22. Teaching blueprints
- 23. FAQ
- 24. Research grounding
- 25. eduKate vocabulary ecosystem
1. What is lexical diversity? The short answer
Lexical diversity is the degree of vocabulary variation in a sample of language. If a 100-word essay uses only a small set of words repeatedly, its lexical diversity is relatively low. If another 100-word essay uses a larger range of different words, its lexical diversity is relatively high. The idea sounds simple; measuring it fairly is not.
The basic unit of the problem is the distinction between tokens and types. Tokens are all word occurrences. Types are distinct word forms under whatever counting rules the analysis uses. In “the cat chased the dog,” there are five tokens but four types if the repeated the counts once as a type.
Lexical diversity therefore does not ask whether the words are difficult, elegant or accurate. It asks how much variation there is. That narrowness is a strength when the measure is used for the right question and a weakness when it is treated as a general score for language quality.
The best use of lexical diversity is descriptive. It can help us notice repetition, compare controlled samples, study development, or examine differences among genres. It should not become a command to replace every repeated word with a synonym.
2. Types and tokens: the counting idea underneath lexical diversity
Imagine the sentence: “A careful reader reads the question carefully.” Depending on the counting method, the word forms careful and carefully are separate types. reader and reads are separate forms too. A lemmatised analysis might instead group inflected forms under a base lemma in some cases. The result therefore depends partly on what the analyst decides a “different word” is.
Orthographic word-form counting is common because it is straightforward: different written forms count separately. But this means run, runs, running and ran can increase diversity even though they belong to one lexical family. Lemma-based counting answers a slightly different question about lexical choice.
Punctuation, contractions, hyphenation, proper names, numbers and multiword expressions introduce further decisions. Is New York one lexical item or two tokens? Does don’t count as one orthographic token or two grammatical units? Automated tools make these choices consistently, but not all tools choose the same rules.
This is why lexical-diversity results should be reported with the method. A number without tokenisation rules, sample length and measure name can look scientific while remaining hard to interpret.
3. Type-token ratio (TTR): the simplest lexical-diversity measure
The classic type-token ratio is simple: divide the number of distinct word types by the total number of word tokens. If a text contains 100 tokens and 70 types, TTR is 0.70. If it contains 100 tokens and 45 types, TTR is 0.45.
This simplicity makes TTR easy to teach. Students can understand that repeated words increase the token count without necessarily increasing the type count. Researchers can compute it without specialised software. It gives an intuitive first look at lexical repetition.
But TTR is not a stable universal property of a writer. The same writer can produce different TTR values depending on topic, length, genre, prompt and the number of unavoidable repeated terms. A science explanation about photosynthesis may legitimately repeat plant, light, energy and carbon dioxide, while a narrative naturally rotates through more concrete actions and descriptions.
TTR is therefore best treated as a sample statistic under controlled conditions, not as a permanent vocabulary score attached to a person.
4. The text-length problem: why TTR usually falls as samples get longer
Short texts can achieve high TTR very easily because almost every new word has a chance to be unique. As a text becomes longer, common grammatical words and central topic words must recur. The denominator continues increasing, while the supply of completely new types grows more slowly. TTR therefore tends to decrease as text length increases.
This means a 50-word paragraph and a 500-word essay cannot be compared naïvely. The longer sample has more opportunity—and often more communicative need—to repeat words. A lower TTR may reflect length rather than poorer lexical range.
Recent research continues to treat text-length sensitivity as a central measurement problem. A 2026 Cambridge study comparing several lexical-diversity indices still used TTR primarily as a benchmark because of its known limitations and examined interactions between diversity measures and text length.
For classroom use, the practical rule is straightforward: if you compare raw TTR, keep samples similar in length and genre. If lengths vary substantially, use a measure designed to reduce length sensitivity or interpret the result with caution.
5. MATTR: moving-average type-token ratio
MATTR attempts to reduce the text-length problem by sliding a fixed-size window across the text. TTR is calculated for each window, and the values are averaged. Instead of letting one long document create a single length-sensitive ratio, MATTR asks what lexical variation looks like locally across many equal-sized stretches.
The choice of window size still matters. A 20-word window captures very local variation; a 100-word window smooths over more of the text. Researchers therefore report the window length because MATTR values from different settings are not automatically comparable.
MATTR is attractive for learner writing because it is relatively interpretable and can work with shorter samples than some alternatives. A 2026 TESOL Quarterly paper used MATTR to examine lexical diversity by word class, showing how noun diversity and verb diversity can reveal different developmental patterns that an all-words score may hide.
In education, the idea behind MATTR is more important than the acronym: fair comparison often requires controlling the amount of language being compared.
6. MTLD: measuring how long lexical variety can be sustained
Measure of Textual Lexical Diversity (MTLD) approaches the problem differently. Rather than returning one raw type-token ratio, it tracks how long a text can proceed before lexical repetition pushes diversity below a chosen threshold, then combines those segments into a score.
The attraction of MTLD is that it was designed to be less dependent on sample length than plain TTR. It has therefore become common in research on second-language writing and lexical proficiency.
MTLD is not “better” in every situation simply because it is mathematically more complex. It still depends on tokenisation, sample characteristics and implementation. Very short samples can be problematic, and different measures may rank texts differently.
For most parents and students, there is no need to calculate MTLD manually. Teachers and researchers may use software, but the instructional question remains: is the learner relying on a narrow set of words because those are the only accessible options, or repeating strategically because the topic requires stable terminology?
7. HD-D, VocD, Maas and other measures: different tools for the same broad problem
Lexical-diversity research contains many indices because no single formula perfectly separates vocabulary variation from text length, genre and sampling effects. Common alternatives include HD-D, VocD, Maas-type measures, standardised TTR and fixed-length sampling approaches.
The existence of many measures is not evidence that the field has failed. It reflects a real measurement difficulty: vocabulary grows non-linearly as text grows. Different measures make different assumptions about how to normalise that growth.
A classroom should not become a contest over acronyms. If the educational purpose is to help a student notice monotonous repetition, a simple count of repeated nouns and verbs may be more actionable than a sophisticated corpus index. If the purpose is research across hundreds of unequal samples, more robust indices matter.
Choose the measure according to the decision you need to make. Measurement should clarify teaching, not replace it.
8. Lexical diversity vs vocabulary richness
The terms lexical diversity and vocabulary richness are sometimes used loosely as if they mean the same thing. In stricter usage, richness is broader. It can include diversity, sophistication, density, rarity, depth or other properties of lexical use.
A text with many different everyday words can have high diversity without sounding lexically sophisticated. A technical paper can have relatively constrained diversity because key terms must repeat, yet still be lexically rich in domain-specific precision.
This distinction matters in feedback. Telling a student to “make the vocabulary richer” could mean reduce clumsy repetition, choose more precise verbs, vary sentence openings, use better topic-specific terms, or deepen explanation. A lexical-diversity score addresses only one of those possibilities.
Good teaching names the intended dimension instead of hiding several goals inside the vague word “better.”
9. Lexical diversity vs vocabulary size
Vocabulary size concerns how many words or word families a person knows. Lexical diversity concerns how varied the vocabulary is in one produced sample. The first is a property of the learner’s broader repertoire; the second is a property of language output under particular conditions.
A learner can have a large vocabulary but deliberately repeat terminology for clarity. Another learner can create high surface diversity in a short essay by cycling through many simple words. The two measures therefore answer different questions.
eduKateSG’s dedicated Vocabulary Size article owns the question of how many words a learner may know or need. This page owns variation in output.
Confusing size with diversity leads to faulty conclusions such as “this writer used 120 unique words, therefore the writer knows only 120 words” or “this student knows 8,000 words, therefore every essay should show high lexical variation.” Neither follows.
10. Lexical diversity vs vocabulary depth
Vocabulary depth concerns how well words are known: meaning precision, multiple senses, morphology, collocation, grammar, register, semantic relationships and other dimensions. Diversity does not measure these directly.
A student can use twenty different words incorrectly and achieve more diversity than another student who uses fifteen words with excellent precision. The first text is more varied; the second may be linguistically stronger.
Depth becomes visible when learners distinguish near-synonyms, handle polysemy, select natural collocations and adapt register. These are qualitative capabilities. They may influence diversity indirectly, because deeper knowledge creates more choices, but one is not a substitute for the other.
This is why vocabulary development should expand both the repertoire and control over that repertoire.
11. Lexical diversity vs lexical sophistication
Lexical sophistication usually concerns how advanced, infrequent, precise or contextually appropriate the words in a sample are, depending on the research framework. Diversity concerns variety.
A sentence can be diverse but unsophisticated: “The dog ran, jumped, barked, rolled and played.” Many distinct content words appear, but they are common. Another sentence can use fewer distinct words while deploying technical terms accurately.
Sophistication itself should not be confused with rarity. An uncommon word is not automatically more appropriate. Modern research increasingly treats lexical sophistication as a multidimensional construct involving frequency, contextual fit and other properties.
For school writing, the operational principle remains: reward precise lexical choices that carry meaning. Do not reward difficulty in isolation.
12. Lexical diversity vs lexical density
Lexical density is another neighbouring concept. It concerns the proportion of content words—typically nouns, lexical verbs, adjectives and many adverbs—relative to grammatical or function words. Diversity concerns how many different words occur.
A dense academic sentence may contain many content words but repeat the same technical vocabulary across the paragraph. A story may use a broad range of verbs, adjectives and nouns while still containing many grammatical words. The two dimensions can move independently.
Lexical density is often related to information packing. Excessive density can make text difficult to read, while low density can make a passage conversational or diffuse. Again, higher is not automatically better.
When evaluating writing, diversity asks about variation; density asks about information-bearing word concentration.
13. Repetition is not automatically a vocabulary weakness
Writers are often taught to avoid repetition, but good prose repeats strategically. A scientific explanation may need to repeat the name of a variable. An argument may repeat a central term so the reader can track the claim. Legal and technical writing often prefers terminological stability over stylistic variety.
Forced synonym substitution can damage cohesion. If a paragraph alternates unpredictably among student, learner, pupil, candidate and child, the reader may wonder whether the terms refer to the same group.
The right question is not “Did this word repeat?” but “Did the repetition help clarity, or did it occur because the writer lacked alternatives?” Functional repetition and accidental repetition look identical to a simple type count.
This is one of the deepest limitations of lexical-diversity metrics: they count forms, not communicative reasons.
14. Lexical diversity in writing
Writing gives learners time to search memory, revise and consult resources, so lexical diversity can look higher in writing than in speech. Writers can replace repeated verbs, reorganise sentences and introduce more specific nouns after the first draft.
Yet writing quality is not maximised by maximising diversity. Cohesion requires some recurrence. Topic control requires stable reference. Key technical terms often should not be varied.
A useful revision approach targets unproductive repetition: vague verbs such as do, get and make when a more specific verb would improve meaning; repeated evaluative adjectives such as good and bad; or repeated sentence frames that make the lexical pattern monotonous.
Measure after meaning. First ask whether the essay communicates clearly. Then inspect lexical range as one diagnostic lens.
15. Lexical diversity in speaking
Speech is produced under time pressure. The speaker must retrieve words while planning grammar, monitoring the listener and managing turn-taking. Repetition can therefore reflect retrieval efficiency, discourse style, hesitation or a deliberate strategy for clarity.
A speaker with a large receptive vocabulary may still use a smaller active set in spontaneous conversation. This is not surprising. Productive vocabulary depends on fast access, not just stored knowledge.
Lexical-diversity analysis of speech must also decide what to do with filled pauses, false starts, repetitions, contractions and discourse markers. Different transcription rules can change the score.
When helping students speak, the goal should be flexible access to useful words and phrases. Numerical variation matters only if it reflects greater communicative control.
16. Genre, topic and task change lexical diversity
A personal narrative, laboratory report, persuasive essay and oral interview place different demands on vocabulary. Comparing their raw diversity scores can therefore confuse genre effects with learner ability.
Narratives often contain varied action verbs, descriptions and concrete nouns. Academic explanations may concentrate around a smaller set of abstract technical terms. Persuasive writing may repeatedly name the central issue while varying evaluative and causal language.
Topic familiarity matters too. A student who knows football deeply can produce diverse, precise vocabulary on that topic but appear lexically narrow when discussing architecture. The underlying vocabulary repertoire is partly domain-sensitive.
Fair assessment therefore controls prompt, genre, time and sample length as much as possible.
17. What lexical diversity can tell teachers—and what it cannot
Lexical diversity can flag a pattern worth investigating. If a student’s essays repeatedly rely on the same small set of general verbs and adjectives even when topics change, there may be a productive vocabulary bottleneck. If diversity rises after extensive reading and vocabulary instruction, that may be one sign of expanding lexical choice.
But the metric cannot tell the teacher why the pattern exists. The cause could be limited vocabulary size, weak retrieval, fear of making mistakes, narrow topic knowledge, exam timing, genre constraints or deliberate clarity.
The teacher therefore needs a diagnostic follow-up: ask the learner to paraphrase, supply alternatives orally, explain why one word fits better than another, and use target vocabulary in new contexts.
A score should start a conversation, not finish the diagnosis.
18. Lexical-diversity assessment: the main cautions
The first caution is length. Plain TTR should not be compared casually across very different sample sizes. The second is genre. The third is tokenisation. The fourth is topic. The fifth is opportunity: a prompt may simply not require varied vocabulary.
The sixth caution is quality. Diversity does not score accuracy, natural collocation, register or semantic precision. A text full of inappropriate synonyms may receive a superficially high variety score.
The seventh caution is interpretation across languages and ages. Morphologically rich languages can create many surface forms. Developmental changes in grammar can alter type counts. Bilingual writers may show different patterns depending on language dominance and topic.
Any high-stakes use therefore requires a validated procedure and multiple measures of language performance.
19. How to improve lexical diversity without forcing synonyms
The safest route to greater lexical variety is not a thesaurus. It is broader knowledge plus stronger retrieval. Read across topics, notice high-utility verbs and nouns, record collocations, practise paraphrase and build semantic networks.
In writing, target repeated functions rather than repeated words. If every paragraph uses “shows,” build a small evidence-verb set: indicates, demonstrates, suggests, reveals, illustrates. Then teach the meaning and evidence strength of each instead of treating them as interchangeable.
In speaking, practise retrieval from scenarios rather than lists. A word becomes useful when it arrives quickly enough for real conversation.
Most importantly, preserve necessary repetition. The goal is flexible choice where choice improves meaning.
A simple TTR calculation
| Text | Tokens | Types | TTR | What it suggests |
|---|---|---|---|---|
| Sample A | 20 | 16 | 0.80 | High surface variety for a very short sample |
| Sample B | 100 | 60 | 0.60 | Moderate variety, but length makes direct comparison with A unfair |
| Sample C | 100 | 75 | 0.75 | More distinct forms than B under the same length condition |
| Sample D | 500 | 210 | 0.42 | Cannot be judged against 100-word samples without accounting for length |
The table illustrates why equal-length samples are so important for raw TTR. Sample D has far more distinct words than any other sample, yet the ratio is lower because a long text must repeat many words. Saying D has “worse vocabulary” would be a measurement error.
20. Worked samples: when greater lexical diversity helps—and when it does not
Repetitive narrative
Before: The boy went to the park. The boy went to the pond. The boy went to the shop.
Diagnosis: The obvious repetition is the boy went. Some repetition is referential, but the repeated general verb suggests limited action vocabulary.
Possible revision: The boy hurried to the park, wandered beside the pond, then stopped at the shop.
What this teaches: Variety improves because the verbs add distinct movement meanings, not because synonyms were inserted randomly.
Technical explanation
Before: Evaporation happens when water gains energy. Evaporation can happen below boiling point. Evaporation occurs at the surface.
Diagnosis: The repeated technical term lowers diversity but strengthens terminological continuity.
Possible revision: Keep evaporation where the concept must remain explicit; vary the surrounding verbs only when meaning improves.
What this teaches: A low diversity score here may be appropriate.
Argument
Before: This policy is bad because the result is bad and the effects are bad.
Diagnosis: The problem is not merely repetition; bad is semantically empty.
Possible revision: This policy is ineffective because it raises costs without improving access, and its long-term effects may be harmful.
What this teaches: The repair increases precision and diversity together.
Description
Before: The room was nice. The curtains were nice. The light was nice.
Diagnosis: Repetition signals a narrow evaluative vocabulary.
Possible revision: The room felt welcoming; soft curtains framed the windows, and warm light spread across the floor.
What this teaches: The writer replaces vague evaluation with concrete information.
Evidence paragraph
Before: The graph shows a rise. It shows a fall later. It shows a stable period.
Diagnosis: Repeated shows may be harmless, but the writer can choose verbs that encode the relationship more precisely.
Possible revision: The graph indicates an early rise, records a later decline, and then levels into a stable period.
What this teaches: Verb variation is useful because each verb fits a different analytical function.
Dialogue
Before: I think the plan is good. I think the timing is good. I think the location is good.
Diagnosis: Conversation naturally repeats frames, so the diversity issue is secondary.
Possible revision: The plan looks workable. The timing suits us, and the location is convenient.
What this teaches: The repair sounds less mechanical while preserving conversational register.
Definition
Before: A habitat is a place where an animal lives. The habitat gives the animal food. The habitat gives the animal shelter.
Diagnosis: Repeated technical term is necessary; repeated gives is less necessary.
Possible revision: A habitat is the environment in which an organism lives, providing access to food, shelter and other resources.
What this teaches: Technical precision improves without excessive synonym swapping.
Report
Before: The survey had 200 people. The survey had ten questions. The survey had two sections.
Diagnosis: The noun repeats because it remains the topic; the verb frame causes monotony.
Possible revision: The survey involved 200 participants, contained ten questions and was organised into two sections.
What this teaches: Verb diversity adds information about different relationships.
Personal reflection
Before: I learned many things. I learned about teamwork. I learned about planning. I learned about patience.
Diagnosis: The repeated frame can be stylistically effective once, but four identical clauses flatten the reflection.
Possible revision: The project taught me how teams coordinate, why planning matters and how patience protects decisions under pressure.
What this teaches: Variation comes from conceptual integration.
Science comparison
Before: A solid has particles close together. A liquid has particles close together. A gas has particles far apart.
Diagnosis: Some repetition supports parallel comparison.
Possible revision: In a solid, particles are closely packed; in a liquid they remain close but can move past one another; in a gas they are much farther apart.
What this teaches: Parallel structure is preserved while vocabulary becomes more precise.
History causation
Before: The war happened because of many reasons. One reason was competition. Another reason was alliances.
Diagnosis: The repeated reason is acceptable but generic.
Possible revision: The war resulted from several interacting causes, including imperial competition and alliance commitments.
What this teaches: The repair introduces disciplinary causal language.
Mathematics explanation
Before: I got the answer by doing the formula. Then I did the numbers. Then I did the answer.
Diagnosis: The general verb do hides mathematical operations.
Possible revision: I substituted the values into the formula, simplified the expression and calculated the final value.
What this teaches: Precise process verbs improve both diversity and explanation.
Literary analysis
Before: The writer uses darkness to show sadness. The writer uses rain to show sadness.
Diagnosis: The repeated frame obscures the relation between devices and effects.
Possible revision: Darkness establishes a bleak atmosphere, while the rain reinforces the character’s isolation.
What this teaches: Lexical variation follows analytical distinctions.
Before: I am writing to ask about the trip. I am writing to ask about the timing. I am writing to ask about payment.
Diagnosis: The repeated opening is unnecessarily formulaic.
Possible revision: I am writing to ask about the trip, particularly the departure time and payment arrangements.
What this teaches: Compression is better than synonym replacement.
Oral explanation
Before: The machine has a part that moves and a part that stops the moving part.
Diagnosis: Low diversity reflects missing technical vocabulary.
Possible revision: The mechanism contains a rotating shaft and a brake that stops its motion.
What this teaches: Learning the domain terms increases both precision and efficiency.
Exam answer
Before: The character is angry because he is angry about the result.
Diagnosis: Repetition exposes circular reasoning rather than only lexical weakness.
Possible revision: The character is furious because the result confirms that his effort has been ignored.
What this teaches: The solution adds causal content and a more precise emotion word.
Summary
Before: The article talks about pollution. It talks about cars. It talks about factories.
Diagnosis: Generic reporting verb repeated.
Possible revision: The article examines pollution from transport and industrial sources.
What this teaches: A more precise reporting verb also allows compression.
Process description
Before: First you put the water in. Then you put the powder in. Then you put the spoon in.
Diagnosis: Repeated put signals missing action verbs.
Possible revision: Pour in the water, add the powder and stir with a spoon.
What this teaches: Verb choice follows the physical action.
Balanced argument
Before: Some people think phones are good. Other people think phones are bad.
Diagnosis: The vocabulary encodes only binary evaluation.
Possible revision: Supporters argue that phones improve access to information, while critics warn that constant notifications can fragment attention.
What this teaches: Lexical variety improves because the argument gains conceptual detail.
Conclusion
Before: In conclusion, I conclude that the conclusion is clear.
Diagnosis: High repetition is stylistically clumsy and conceptually empty.
Possible revision: Overall, the evidence supports the view that the policy should be revised.
What this teaches: A clearer claim removes redundant lexical material.
21. Diagnostic atlas: 25 lexical-diversity patterns
High TTR, weak writing
What it may mean: The writer constantly changes words but many choices are imprecise or unnatural. Better next move: Prioritise meaning, collocation and register. Do not reward the score itself. Why: Surface variation can hide shallow word knowledge.
Low TTR, strong technical writing
What it may mean: The text repeats key terminology accurately. Better next move: Check whether repetition is functional before changing anything. Why: Domain clarity may require stable labels.
Low diversity across every genre
What it may mean: The learner relies on a narrow productive vocabulary. Better next move: Build high-utility verb, adjective and noun networks, then practise retrieval in new contexts. Why: Consistent cross-task repetition is more diagnostic than one low score.
High diversity only in prepared writing
What it may mean: The learner uses resources or long planning time effectively but cannot retrieve the same range in speech. Better next move: Train oral retrieval and timed writing. Why: The gap is access, not necessarily vocabulary size.
Diversity falls in longer essays
What it may mean: Text-length effect may be driving the change. Better next move: Compare equal-length segments or use a length-resistant measure. Why: Do not interpret plain TTR across unequal sample sizes.
Diversity rises after thesaurus use
What it may mean: Many substitutes appear but meaning becomes unstable. Better next move: Teach near-synonym boundaries and collocation. Why: Numerical variety is not communicative quality.
Same adjectives repeat
What it may mean: The learner lacks evaluative precision. Better next move: Build semantic scales: useful/effective/efficient/reliable/appropriate rather than random synonyms. Why: Target the functional category causing monotony.
Same verbs repeat in analysis
What it may mean: The learner overuses show, say, get or make. Better next move: Teach reporting and analytical verbs with evidence-strength differences. Why: Verb diversity often gives the largest practical improvement.
Same nouns repeat
What it may mean: The topic may require them. Better next move: Check referential clarity before replacing nouns with pronouns or synonyms. Why: Noun repetition is not automatically a flaw.
Speech diversity is lower than writing
What it may mean: Real-time retrieval constrains lexical choice. Better next move: Use timed cue-to-word and phrase-retrieval drills. Why: This is a common modality effect.
Bilingual learner uses narrow English vocabulary
What it may mean: The concept may exist strongly in another language. Better next move: Map existing concepts to English forms and collocations. Why: Do not confuse English lexical access with conceptual weakness.
Very high diversity in a short paragraph
What it may mean: Short length inflates TTR. Better next move: Expand the sample or compare fixed-length windows. Why: Small samples can create unstable scores.
Student avoids repeating key term
What it may mean: Synonym variation causes reference ambiguity. Better next move: Teach intentional repetition for cohesion. Why: A reader must know whether two labels refer to the same concept.
Vocabulary looks diverse but generic
What it may mean: Many different common words appear without technical precision. Better next move: Teach domain terms and high-utility academic vocabulary. Why: Diversity and sophistication are separate dimensions.
Vocabulary is sophisticated but repetitive
What it may mean: A small set of advanced terms dominates. Better next move: Decide whether topic concentration explains it; if not, expand functional alternatives. Why: Rarity and diversity measure different properties.
Automated score changes after tokenisation
What it may mean: Hyphens, contractions, names or lemmatisation altered the type count. Better next move: Keep preprocessing rules constant. Why: Measurement choices can create apparent linguistic change.
Two tools give different diversity scores
What it may mean: They may use different formulas or windows. Better next move: Compare like with like and document settings. Why: Metric disagreement is not necessarily an error.
Student’s score rises but comprehension does not
What it may mean: Production variety improved but receptive knowledge may not. Better next move: Assess reading vocabulary and word depth separately. Why: Lexical diversity is not a complete vocabulary test.
Student’s score falls after learning a subject
What it may mean: New technical words repeat frequently. Better next move: Inspect which terms repeat and whether they carry necessary content. Why: Learning can lower surface diversity while increasing conceptual precision.
AI-generated text scores highly
What it may mean: The system can vary vocabulary easily. Better next move: Assess argument, evidence, accuracy and human ownership separately. Why: High lexical diversity is not proof of understanding or authorship.
Narrative score higher than report score
What it may mean: Genre produces different lexical opportunities. Better next move: Compare within genre. Why: Cross-genre scores answer a different question.
Student pads with adjectives
What it may mean: Diversity rises but prose becomes verbose. Better next move: Teach economy and strong nouns/verbs. Why: More types can reduce quality if they add no information.
Student repeats connective words
What it may mean: Function words can dominate token counts. Better next move: Consider content-word diversity or POS-specific measures where appropriate. Why: All-word diversity can hide which lexical class is changing.
Noun diversity improves, verb diversity does not
What it may mean: Development is uneven by word class. Better next move: Target verb networks and event language. Why: POS-specific analysis can reveal useful detail.
Vocabulary diversity plateaus
What it may mean: The learner’s reading environment may be too narrow. Better next move: Broaden genres and subjects, then promote useful encountered words into active practice. Why: Output range depends partly on input range.
22. Ten teaching blueprints for lexical diversity
Repeated-verb audit
Activity: Take a 300-word student essay. Highlight every lexical verb. Count the five most repeated verbs. Decide which repetitions are functional and which are vague. Replace only the vague ones with more precise verbs, then explain how meaning changed. Learning purpose: This teaches students that diversity should follow semantic improvement.
Equal-length TTR demonstration
Activity: Prepare two 100-word passages and one 400-word passage. Calculate simple TTR. Then take a 100-word slice from the long passage and compare again. Learning purpose: Students see the text-length problem directly instead of memorising a warning.
Synonym danger lab
Activity: Give five near-synonyms and ten sentences. Students choose the best fit and justify collocation, strength and tone. Learning purpose: It separates lexical diversity from indiscriminate substitution.
Genre comparison
Activity: Compare a narrative and science explanation written by the same learner. Identify what repetition is required by genre. Learning purpose: Students learn that diversity is task-dependent.
MATTR intuition activity
Activity: Without calculating MATTR formally, divide a long text into equal windows and count unique words in each. Compare local variation. Learning purpose: This makes moving-window logic concrete.
Content-word diversity
Activity: Separate nouns, verbs, adjectives and adverbs from function words in a short text. Examine which class is repetitive. Learning purpose: Teachers can target the lexical category causing monotony.
Paraphrase with constraints
Activity: Rewrite a sentence in three ways while preserving the exact proposition. Ban one vague verb but require the key technical noun to stay unchanged. Learning purpose: The task trains flexible expression without damaging terminological consistency.
Oral-to-written comparison
Activity: Have the learner speak for one minute on a topic and write 150 words on the same topic. Compare repeated lexical items. Learning purpose: The contrast reveals retrieval differences between modalities.
Vocabulary network expansion
Activity: Choose one overused word such as good, bad, important, shows or says. Build a network organised by meaning differences rather than synonym lists. Learning purpose: A structured network creates usable alternatives.
AI score critique
Activity: Generate two passages with noticeably different lexical variety. Ask students which is clearer and why. Then discuss why a numerical diversity score cannot decide quality alone. Learning purpose: This develops metric literacy in an AI-rich environment.
23. Frequently asked questions about lexical diversity
What is lexical diversity in simple words?
Lexical diversity is how varied the words are in a sample of speaking or writing.
What is a type in lexical diversity?
A type is a distinct word form under the counting rules being used.
What is a token?
A token is one occurrence of a word in the language sample, including repeated occurrences.
What is type-token ratio?
TTR is the number of distinct word types divided by the total number of word tokens.
Is a high TTR always good?
No. Short texts inflate TTR, and useful repetition can lower it. Accuracy and clarity matter more than maximising the ratio.
Why does TTR go down in longer texts?
Longer texts inevitably repeat common grammatical and topic words, so tokens grow faster than new types.
What is MATTR?
MATTR is moving-average type-token ratio. It computes TTR across fixed-size moving windows and averages them to reduce length sensitivity.
What is MTLD?
MTLD is a measure designed to estimate how long lexical variety can be sustained before diversity falls below a set threshold.
Which lexical-diversity measure is best?
There is no universal best measure. The choice depends on text length, sample type, research purpose and available validation.
Is lexical diversity the same as vocabulary richness?
Not always. Vocabulary richness is often used more broadly and may include sophistication, density or depth as well as variation.
Is lexical diversity the same as vocabulary size?
No. Size is the learner’s broader repertoire; diversity is the variation visible in one sample.
Is lexical diversity the same as lexical sophistication?
No. Sophistication concerns properties such as rarity, advancement or contextual precision; diversity concerns variety.
Is lexical diversity the same as lexical density?
No. Density is the proportion of content words; diversity is the number or distribution of distinct lexical items.
Can repetition improve writing?
Yes. Repetition can support cohesion, terminological precision and reader tracking when it is purposeful.
Should students replace repeated words with synonyms?
Only when the replacement preserves meaning, collocation, register and reference. Forced synonym replacement often weakens writing.
How can students improve lexical diversity?
Read widely, learn high-utility word networks, practise paraphrase, strengthen retrieval and revise vague repeated verbs or adjectives.
Can lexical diversity measure vocabulary knowledge?
It provides evidence about produced variety, but it does not directly measure total vocabulary size, depth or receptive knowledge.
Does lexical diversity predict language proficiency?
Research often finds associations, especially in controlled learner samples, but the relationship depends on measure, task, language and text length.
Can lexical diversity be used in marking essays?
It can be one analytic feature, but should not replace human judgement of content, organisation, accuracy, coherence and appropriateness.
Does AI have high lexical diversity?
AI systems can readily produce varied vocabulary, but diversity does not demonstrate understanding, factual accuracy or authorship.
Why might technical writing have low lexical diversity?
Specialised writing often repeats core terms intentionally to maintain precision.
Why might narratives have higher lexical diversity?
Narratives can draw on many actions, objects, settings and descriptions, creating more opportunities for varied lexical choice.
Can spoken language have lower diversity than written language?
Yes. Real-time retrieval and planning constraints often reduce the range of words available during spontaneous speech.
What is POS-specific lexical diversity?
It measures diversity separately for word classes such as nouns or verbs, revealing patterns hidden by one overall score.
What is vocabulary variety?
In everyday educational language, vocabulary variety usually means using a sufficiently broad range of appropriate words rather than relying on the same general words repeatedly.
What is a good TTR score?
There is no universal good score. TTR depends heavily on sample length, genre, tokenisation and task.
Can two texts with the same TTR be equally good?
Not necessarily. One can be accurate and coherent while the other uses many inappropriate word choices.
How long should a text be for lexical-diversity analysis?
It depends on the measure. Very short samples are unstable, and researchers choose methods appropriate to sample length.
Should teachers calculate MTLD for every student?
Usually not. The measure can be useful for research or specialised assessment, but classroom teaching often benefits more from targeted analysis of repeated vocabulary and word choice.
What is the main lesson of lexical diversity?
Variety is useful when it reflects flexible, precise lexical choice. A diversity number should never become the goal by itself.
24. Research grounding
Current and established research agrees on two broad points. First, lexical diversity is a meaningful feature of language production and can relate to proficiency. Second, measurement is sensitive to sample length and method. A 2026 Cambridge study compared several indices—including TTR, MATTR and MTLD variants—and still found interactions with text length. A 2026 TESOL Quarterly study introduced part-of-speech-specific MATTR measures to make lexical-development patterns more interpretable. Earlier work and current reviews continue to warn that plain TTR is especially sensitive to text length.
- Cambridge Core, 2026: multiple lexical-diversity measures and text-length sensitivity.
- TESOL Quarterly, 2026: part-of-speech-specific lexical diversity using MATTR.
- Applied Linguistics, 2024: challenges in comparing lexical diversity across texts of different lengths.
- Frontiers in Psychology: lexical diversity, sophistication and the limitations of simple TTR.
- The Modern Language Journal: diversity, sophistication and fluency as distinct dimensions of lexical proficiency.
25. The eduKate vocabulary ecosystem
Lexical diversity is one small instrument inside a much larger vocabulary system. The pages below own neighbouring questions so this article does not cannibalise them.
- What Is Vocabulary? — canonical definition, types, examples and importance.
- Vocabulary — master router from Primary 1 to adult and career vocabulary.
- Vocabulary Learning Hub — level-based learning routes.
- Vocabulary Size — how many words a learner knows or needs.
- Word Knowledge — what it means to know a word deeply.
- Mental Lexicon — storage, connection and retrieval.
- Vocabulary Breadth and Depth — expansion and deepening across learning phases.
- Lexical Coverage — how much vocabulary is needed to understand text.
- Expressive Vocabulary — words available for independent use.
- Receptive Vocabulary — words understood in reading and listening.
The final answer
Lexical diversity is the variety of words visible in a sample of language. It is useful because repetition patterns can reveal something about lexical choice, development and style. It is limited because the number changes with text length, genre, topic, tokenisation and communicative purpose.
Type-token ratio gives the simplest introduction: distinct types divided by total tokens. Measures such as MATTR and MTLD attempt to reduce some of the length problem. None of them turns lexical diversity into a universal score for writing quality.
The educational goal is therefore not to maximise variety. It is to build enough vocabulary knowledge and retrieval flexibility that a learner can choose a different word when a different word improves meaning, while repeating the same word when repetition protects clarity, cohesion or technical precision.
Good vocabulary is not maximum variation. It is controlled choice.
Appendix A. Precision bank: 20 overused words and how to diversify them safely
A common route to greater lexical diversity is to examine a few high-frequency vague words. The purpose is not to ban them. Common words are essential. The purpose is to recognise when a more specific word would carry information that the general word leaves unstated.
good
Possible alternatives by meaning: effective, useful, beneficial, strong, appropriate, convincing, reliable. Caution: Do not treat these as synonyms. Effective asks whether something works; efficient would add resource use; convincing concerns persuasion; reliable concerns consistency. Example: The programme was effective because completion rates rose, but it was not efficient because staffing costs doubled.
bad
Possible alternatives by meaning: harmful, ineffective, weak, inaccurate, inappropriate, unreliable, damaging. Caution: First name the failure. “Bad” hides whether the problem is morality, performance, evidence, accuracy or consequences. Example: The method was unreliable because identical trials produced widely different results.
big
Possible alternatives by meaning: large, major, substantial, significant, extensive, severe, considerable. Caution: Choose according to what is large: physical size, effect, quantity, scope or seriousness. Example: The policy produced a substantial reduction in waiting time.
small
Possible alternatives by meaning: minor, slight, limited, narrow, modest, negligible, compact. Caution: A small physical object is different from a minor effect or negligible difference. Example: The improvement was modest and did not justify the additional cost.
show
Possible alternatives by meaning: indicate, demonstrate, reveal, illustrate, suggest, record, display. Caution: Reporting verbs carry different evidence strength. Suggest is cautious; demonstrate is stronger. Example: The pattern suggests a relationship, but the study does not demonstrate causation.
say
Possible alternatives by meaning: state, argue, claim, observe, note, explain, suggest, insist. Caution: The right reporting verb depends on what the source is doing. Example: The author argues that public transport should be treated as social infrastructure.
get
Possible alternatives by meaning: obtain, receive, become, understand, reach, acquire, retrieve. Caution: The correct replacement depends entirely on the intended meaning of get. Example: Students acquire vocabulary through repeated encounters but must retrieve it during production.
make
Possible alternatives by meaning: create, produce, construct, cause, form, generate, establish. Caution: Many fixed collocations still require make, such as make a decision. Diversity should not destroy conventional phrases. Example: The new rule generated confusion because its conditions were not explicit.
thing
Possible alternatives by meaning: factor, object, issue, feature, element, event, process, condition. Caution: Identify the category instead of reaching for a random sophisticated noun. Example: One factor affecting recall was the delay between study and testing.
important
Possible alternatives by meaning: central, significant, critical, relevant, influential, necessary, consequential. Caution: Importance can mean centrality, necessity, effect or relevance. Name the relationship. Example: Accurate vocabulary is critical when two technical terms mark different concepts.
very
Possible alternatives by meaning: highly, deeply, strongly, extremely, particularly—or remove it. Caution: Many adjective-adverb pairings are collocational, not interchangeable. Example: The result was highly significant statistically but only slightly important practically.
nice
Possible alternatives by meaning: pleasant, welcoming, considerate, attractive, enjoyable, elegant. Caution: Replace vague approval with the quality being evaluated. Example: The room felt welcoming because natural light and quiet colours softened the space.
interesting
Possible alternatives by meaning: surprising, revealing, unusual, significant, puzzling, engaging. Caution: Say what kind of interest the information creates. Example: The finding is revealing because it reverses the pattern seen in earlier studies.
problem
Possible alternatives by meaning: difficulty, limitation, obstacle, error, risk, conflict, constraint. Caution: Each alternative identifies a different structure. Example: The small sample is a limitation, not an error in the analysis.
help
Possible alternatives by meaning: support, enable, facilitate, improve, reduce, strengthen, assist. Caution: Choose the mechanism rather than the generic benefit. Example: Morphology can support inference by revealing meaningful word parts.
change
Possible alternatives by meaning: increase, decrease, shift, fluctuate, transform, revise, modify. Caution: Name direction, scale or mechanism when possible. Example: Vocabulary access improved after spaced retrieval practice.
use
Possible alternatives by meaning: apply, employ, deploy, operate, consume, draw on. Caution: Many contexts are still best served by the common verb use; variation is optional. Example: Writers deploy technical terms when the distinction they carry matters.
look
Possible alternatives by meaning: appear, examine, inspect, observe, glance, seem. Caution: Some alternatives change from physical seeing to inference or appearance. Example: The results appear stable, but closer inspection reveals a subgroup effect.
think
Possible alternatives by meaning: believe, infer, reason, conclude, assume, judge, consider. Caution: Choose the cognitive operation. Example: From the evidence, we can infer that the treatment changed response speed.
really
Possible alternatives by meaning: genuinely, substantially, actually, clearly—or remove it. Caution: Intensifiers often add little. Replace only if a specific relationship is intended. Example: The effect was substantial, not merely statistically detectable.
Appendix B. Thirty classroom and writing scenarios
School report
A student writes “The experiment was good.” Ask: good in what way—accurate, reliable, efficient, safe, well-controlled or informative? Rewrite only after the criterion is clear. Principle: Lexical diversity grows out of judgement, not decoration.
Narrative action
A story repeats “went” eight times. Sort the movements: hurried, wandered, climbed, returned, approached, escaped, entered, crossed. Use only verbs that match the physical action. Principle: Verb variety can simultaneously improve imagery and precision.
Argument evidence
An essay repeats “shows.” Classify each source relationship: demonstrates, suggests, indicates, illustrates, reports, reveals. Use stronger verbs only when the evidence warrants them. Principle: Reporting verbs encode epistemic strength.
Science terminology
A student wants synonyms for “cell” to avoid repetition. Stop the substitution. The technical term should remain stable unless another term names a genuinely different concept. Principle: Some repetition protects scientific accuracy.
Mathematics explanation
A learner writes “do the number, do the formula, do the answer.” Replace with operations: substitute, simplify, multiply, divide, solve, calculate, verify. Principle: Vocabulary diversity can reveal mathematical reasoning.
Oral presentation
A speaker repeats “important” because retrieval is slow. Prepare three meaning-specific alternatives and rehearse them in complete phrases, not isolated cards. Principle: Phrase retrieval is faster than searching for single words under pressure.
Summary writing
A student varies every occurrence of “the study” with research, investigation, paper, experiment and article even though some labels are inaccurate. Restore stable reference. Principle: Cohesion outranks diversity.
Literary response
A paragraph repeats “sad.” Build an emotion gradient: disappointed, lonely, despondent, grief-stricken, resentful. Match each to textual evidence. Principle: Variation becomes interpretation when distinctions are evidence-based.
History essay
The writer repeats “cause.” Add trigger, contributing factor, structural condition and catalyst only after teaching their causal roles. Principle: Near-synonyms can encode different positions in a causal model.
Geography answer
The word “increase” repeats. Decide whether values rose gradually, surged, climbed, doubled, fluctuated upward or expanded. Principle: Data description benefits from controlled process vocabulary.
Email writing
A formal email repeats “I want.” Replace with purpose-specific forms: I would like to request, I am writing to enquire, Could you confirm, I would appreciate. Principle: Register and function shape lexical choice.
Primary composition
A learner repeats “nice” and “good.” Instead of supplying advanced adjectives, ask for sensory or behavioural evidence. Principle: Concrete detail often solves repetition better than synonym lists.
Debate speech
A speaker repeats “bad for society.” Break the claim into costs, risks, inequity, reduced access or long-term harm. Principle: Better vocabulary follows better argument structure.
Reading response
A student uses diverse vocabulary copied from the passage but cannot explain it. Ask for paraphrase without looking. Principle: Textual diversity is not the same as owned vocabulary.
Bilingual task
A learner knows a precise concept in another language but uses a generic English word. Teach the English lexical mapping and common collocations. Principle: Concept knowledge can support rapid lexical expansion.
AI-assisted essay
AI replaces repeated words with rare alternatives. The student audits every replacement for meaning, register and collocation, reverting those that distort the claim. Principle: Lexical variety generated automatically still requires human judgement.
Research abstract
A writer worries that “participants” repeats. Keep it if alternative labels would imply different groups. Principle: Research writing values referential stability.
Procedure writing
A method repeats “then.” Rather than substitute connectors mechanically, restructure steps with imperative verbs and grouping. Principle: Syntactic redesign can reduce repetition without lexical ornament.
Speech transcript
Filled pauses and discourse markers dominate tokens. Decide whether the diversity measure should include them for the research question. Principle: Preprocessing rules must match the construct.
Portfolio review
Compare five essays of equal length from the same student across a term. Look for changes in repeated content words, not one isolated score. Principle: Longitudinal patterns are more meaningful when tasks are comparable.
Vocabulary lesson
Students receive ten synonyms for “happy.” Group them by intensity, cause and register before any writing task. Principle: Unstructured lists create apparent choice without usable distinctions.
Exam revision
A student memorises rare adjectives. Replace half the list with high-utility analytical verbs and academic nouns. Principle: Useful lexical diversity comes from functions the exam actually demands.
Spoken explanation
The learner can name many alternatives on paper but repeats one word orally. Practise three-second retrieval from varied cues. Principle: Access speed is the bottleneck.
Group discussion
Participants start using the same terms. This lexical alignment may improve shared understanding, even though diversity falls. Principle: Lower diversity can sometimes signal successful coordination.
Technical documentation
The text intentionally repeats exact interface labels. Do not vary them stylistically. Principle: User safety and clarity can require repetition.
Creative writing
A passage repeats a key word for rhetorical effect. Keep the repetition if rhythm or thematic emphasis depends on it. Principle: Metrics cannot automatically recognise artistic purpose.
Young learner speech
Short samples produce unstable ratios. Collect several comparable samples before concluding that vocabulary variety changed. Principle: Developmental interpretation needs adequate sampling.
Assessment rubric
A rubric says “wide vocabulary.” Translate this into observable descriptors: precise word choice, appropriate register, varied verbs where useful, stable technical terms and controlled repetition. Principle: Rubric language becomes teachable when operationalised.
Parent feedback
Instead of “use more vocabulary,” say “You rely on good, bad and nice; this week we will build words for effectiveness, emotion and appearance.” Principle: Specific feedback creates a tractable learning target.
Teacher planning
A class has low lexical variety in explanations. Add oral rehearsal of subject verbs before writing. Principle: Instruction should target the actual lexical function missing.
Appendix C. Thirty principles of lexical diversity
Variety follows knowledge
A writer cannot reliably vary vocabulary that is not available in memory. Long-term lexical diversity grows through reading, listening, study, discussion and retrieval. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Variation needs semantic control
The ability to produce alternatives matters only when the learner understands how they differ. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Repetition has functions
Cohesion, emphasis, technical precision and shared reference can all require repeated words. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Length changes the metric
Raw type-token ratio should not be compared casually across samples of different length. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Genre shapes opportunity
Stories, reports, explanations and conversations create different lexical distributions. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Topic knowledge shapes output
People produce more precise and varied vocabulary in domains they know well. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Speech and writing differ
Writing permits revision and lexical search; speech requires real-time access. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Tokens need rules
Different tokenisation and lemmatisation procedures can change results. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
One number is not vocabulary
Diversity does not measure total vocabulary size, receptive knowledge, depth, collocation or accuracy. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
High scores can mislead
Random synonym substitution can increase variety while degrading meaning. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Low scores can mislead
Technical clarity can require repeated terminology. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Word class matters
Noun diversity, verb diversity and adjective diversity can develop differently. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Comparison needs control
Equal length, comparable prompts and similar genres improve interpretability. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Longitudinal evidence is stronger
Repeated comparable samples can reveal development more reliably than one essay. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Revision should start with meaning
Identify vague or repeated functions first, then choose better words. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Common words remain essential
High-frequency vocabulary carries grammar, cohesion and everyday meaning. The aim is not to eliminate it. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Precision can reduce diversity
Choosing one exact technical term repeatedly may be better than rotating through approximate labels. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Paraphrase tests flexibility
Expressing the same proposition in several accurate ways reveals controlled lexical choice. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Retrieval practice supports production
Words that can be retrieved from multiple cues are more likely to appear in authentic writing and speech. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Reading expands the candidate pool
Wide reading supplies lexical alternatives and collocational patterns that deliberate lists cannot provide alone. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Collocation constrains choice
A synonym that fits the definition may not fit the phrase. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Register constrains choice
Formal, conversational, technical and literary settings privilege different vocabulary. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Assessment should be multidimensional
Combine diversity with accuracy, sophistication, depth, coherence and task achievement. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
AI makes variety cheap
Because tools can vary wording instantly, human judgement about meaning becomes even more important. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Teaching should preserve agency
Students should learn why a word is better, not merely accept automated substitutions. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Metrics should serve decisions
Use lexical diversity only when the result informs a real educational or research question. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Do not optimise the proxy
Once students chase TTR directly, the metric can stop reflecting genuine vocabulary growth. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Vocabulary is relational
Words gain usefulness through networks of meaning, morphology, collocation and context. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Transfer validates learning
A varied word used correctly in a new task is stronger evidence than one inserted from a visible word bank. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Controlled choice is the goal
The mature writer can vary when variety clarifies and repeat when repetition clarifies. A useful classroom question is: does changing this word improve the meaning, or only change the surface? That question keeps lexical diversity connected to communication rather than turning it into a numerical target.
Advanced Masterclass: How to Read Lexical Diversity Without Being Misled
Lexical diversity becomes most useful when the reader can separate the number produced by a measure from the linguistic story that produced the number. The next sections treat the score as evidence rather than verdict. They show how length, genre, repetition, morphology, topic, word class, revision, speaking pressure, bilingualism and artificial intelligence can change lexical-diversity values without producing a simple corresponding change in language quality.
26. Sampling: a lexical-diversity score belongs to a sample, not to a person
A student does not possess one permanent lexical-diversity score in the way a ruler possesses a fixed length. The score is generated by a particular sample of language under particular conditions. Change the prompt, the genre, the time limit or the amount of text, and the score can change even when the writer’s underlying vocabulary knowledge has not.
This is a basic measurement principle. A sample is evidence about a repertoire, not the repertoire itself. A learner may produce narrow vocabulary in an unfamiliar topic because the relevant concepts are missing, then produce highly varied and precise language when discussing a domain of expertise.
Good comparison therefore requires repeated samples. One essay can suggest a pattern; several comparable essays can establish whether the pattern is stable. If lexical diversity is low across narratives, explanations, arguments and oral responses, the case for a productive vocabulary bottleneck becomes stronger.
For classroom diagnosis, collect small but comparable samples over time. Keep length, genre and prompt difficulty reasonably stable, then examine whether the same vague verbs and adjectives continue to dominate.
27. Window size in MATTR: a parameter is part of the result
MATTR sounds objective, but the analyst still chooses a window size. A small window examines very local variation. A larger window measures diversity over a broader stretch of language. The same text can therefore receive different MATTR values under different window settings.
This does not make MATTR useless. It means the parameter must be reported and comparisons must use the same settings. A result described only as “MATTR = 0.72” is incomplete if the reader does not know the window size and tokenisation procedure.
In student writing, a smaller window may be useful when samples are short. In long texts, larger windows can smooth over local bursts of repetition. Researchers choose according to the question and available sample length.
The educational lesson is wider than MATTR: every metric contains design choices. Numbers become trustworthy when the choices are visible and consistent.
28. Word forms, lemmas and families: what counts as a different word?
Suppose a student writes analyse, analyses, analysed, analysing, analysis, analytical. A surface-form measure may count several of these as different types. A lemma-based analysis may group some forms. A word-family analysis may connect even more of them.
Each method answers a different question. Surface forms capture morphological variation visible in the text. Lemmas focus more closely on lexical choices after inflection. Word families reflect broader morphological knowledge but require stronger assumptions about what learners know.
This matters when comparing languages. Languages vary in how much grammatical information is expressed through word endings. A form-based metric can therefore reflect morphology as well as lexical choice.
When teachers talk about “unique words,” they should remember that the definition of unique depends on the counting method.
29. Function words and content words tell different stories
All-word lexical diversity includes articles, prepositions, pronouns, auxiliaries and other high-frequency grammatical words. These items repeat heavily because grammar requires them. Their repetition can lower diversity even when content vocabulary is broad.
Content-word diversity looks instead at lexical nouns, verbs, adjectives and many adverbs. This can be more directly relevant to questions about vocabulary range, though the exact classification still depends on the analysis.
Part-of-speech-specific measures go further. A learner may show strong noun diversity because topic knowledge is rich but weak verb diversity because explanations rely on do, make, get and show. One overall score can hide this.
The 2026 TESOL Quarterly work on POS-specific lexical diversity is useful precisely because it makes these different developmental patterns more interpretable.
30. Cohesion can require repetition
A coherent text helps the reader track entities and ideas. Repeating the same key noun is often the safest way to maintain reference. Replacing that noun with several near-synonyms can make the prose look more varied while making the logic harder to follow.
Academic writing particularly values terminological stability. If a paper defines lexical diversity, alternating with vocabulary richness, word variety and lexical richness may imply that these are exact equivalents when they may not be.
Rhetorical repetition can also create emphasis and rhythm. Speeches, narratives and persuasive writing deliberately repeat key phrases to make a pattern memorable.
A lexical-diversity metric sees recurrence. A human reader must decide whether the recurrence is a weakness, a cohesion device or a rhetorical choice.
31. Revision changes lexical diversity even without vocabulary growth
A learner can raise lexical diversity during revision by replacing repeated words, combining sentences or deleting redundant phrases. The student’s vocabulary repertoire may be unchanged; what changed is the ability to edit the sample.
This distinction matters when comparing timed and untimed writing. An untimed essay benefits from searching, rereading and external resources. A timed exam samples more immediate productive access.
Both performances are valid for different purposes. Revision demonstrates editorial control. Timed writing demonstrates retrieval under constraint. A portfolio and an exam should not be interpreted as if they measure the same lexical capability.
If a teaching intervention claims to improve vocabulary knowledge, assess delayed and unassisted production as well as edited final work.
32. Topic knowledge can raise or lower diversity
People possess richer lexical networks in domains they know well. A student who follows astronomy may readily produce orbit, atmosphere, telescope, radiation, gravity, eclipse, satellite and trajectory. On an unfamiliar topic, the same student may rely on general nouns and verbs.
Topic knowledge supplies more than terminology. It supplies distinctions. The learner knows which properties matter, which processes exist and which relationships require naming. Vocabulary and conceptual knowledge therefore reinforce each other.
A high lexical-diversity score on a familiar topic can partly reflect knowledge rather than general language proficiency. A low score on an unfamiliar topic can partly reflect conceptual uncertainty.
Fair assessment uses several topics or controls topic familiarity when possible.
33. Development: children can become better speakers while TTR falls
As children produce longer, more grammatically complex language samples, simple TTR can decrease even though vocabulary knowledge is growing. Longer samples contain more necessary repetition. This is one reason raw TTR has historically behaved unexpectedly in developmental research.
A young child producing twenty words may use fifteen unique forms and receive a very high ratio. An older child producing two hundred coherent words will inevitably repeat many common items, lowering the ratio.
Developmental interpretation therefore needs age-appropriate measures, sufficient sample length and attention to grammar and discourse. The same numerical scale should not be used carelessly across very different developmental stages.
The key lesson is that language growth is multidimensional. One ratio cannot represent vocabulary, syntax, fluency and discourse development at once.
34. Bilingual lexical diversity is not a deficit score
A bilingual speaker’s vocabulary is distributed across languages, contexts and domains. The person may know family and cultural terms more richly in one language, academic terms more richly in another, and switch depending on audience or topic.
Analysing only one language can therefore underestimate the person’s total conceptual and lexical resources. Code-switching also complicates tokenisation: should words from another language count, and under which lexicon? The answer depends on the research question.
Cross-language cognates can support retrieval, while competing lexical forms can sometimes slow access. Both effects influence observed production.
Educational interpretation should therefore ask what language repertoire the task allows before treating low single-language diversity as evidence of weak thinking or limited knowledge.
35. Goodhart’s Law for vocabulary: when a metric becomes the target
Any measure can be gamed once people optimise for the score rather than the underlying skill. If students are told that a higher TTR is always better, they can inflate it by replacing repeated words with unnecessary alternatives, avoiding useful function words, or shortening the text.
The score rises while communication may worsen. This is a classic proxy problem: lexical diversity is meant to provide evidence about variation, not become a direct objective.
The same risk appears in automated writing systems. A model can be prompted to “increase lexical diversity” and produce more varied vocabulary immediately. That does not prove deeper reasoning, stronger evidence or improved human learning.
Use metrics diagnostically and invisibly where possible. Teach the linguistic behaviour that matters—precision, flexible retrieval, appropriate variation—rather than teaching students to optimise a number.
36. Lexical diversity cannot reliably tell you whether AI wrote a text
AI-generated writing can show high lexical diversity, low lexical diversity, sophisticated vocabulary or deliberate simplicity depending on the prompt and model. Human writing can show the same range. There is no principled reason to treat one diversity value as an authorship detector.
A text may also be collaboratively produced: human ideas, AI drafting, human revision. Lexical statistics cannot reconstruct that process from surface variety alone.
If authorship matters, use process evidence, version history, source discussion, oral defence and transparent classroom policy rather than a lexical-diversity threshold.
The educational use of lexical metrics should stay focused on language analysis, not unsupported forensic claims.
37. A better model: diversity × precision × appropriateness × cohesion
Instead of asking whether vocabulary is “varied enough,” evaluate several dimensions. Diversity asks about range. Precision asks whether the chosen word expresses the intended distinction. Appropriateness asks whether register and collocation fit. Cohesion asks whether lexical choices help the reader track the text.
A strong text does not need maximum scores on every dimension. Technical writing may accept lower diversity for higher precision and cohesion. Creative writing may exploit greater variation for imagery. Conversation may prefer common vocabulary for speed and shared understanding.
This four-part model prevents the common mistake of equating variety with quality. It also creates better feedback because the teacher can name which dimension is weak.
Vocabulary teaching improves when every revision has a reason.
38. Forty extended cases: interpreting lexical diversity in real language
The repeated “important” problem
Situation: A student writes important eight times in 400 words. Analysis: Before replacing it, classify each meaning: necessary, influential, central, relevant, serious, consequential or statistically significant. What to learn: The student discovers that three repetitions are appropriate because the same idea is being maintained, while five conceal different judgements. Those five can be revised with words that encode the intended relationship. The gain is not variety for its own sake; the vocabulary now carries distinctions the original adjective erased.
The repeated technical noun
Situation: A biology explanation uses cell twelve times. Analysis: The writer worries that the repetition will lower lexical diversity and searches for synonyms. What to learn: There may be no good substitute because cell names a precise technical entity. Pronouns can reduce some repetition where reference remains clear, but replacing the term with vague alternatives such as unit or structure can damage accuracy. The correct decision may be to accept lower diversity.
The synonym chain
Situation: An essay alternates student, learner, pupil, candidate, child. Analysis: The writer believes repeated nouns are always weak style. What to learn: The terms are not exact equivalents. Candidate implies participation in an assessment or selection; child marks age; pupil varies by dialect and context. If one population is being discussed, a stable label is usually more coherent.
The strong-verb revision
Situation: A narrative repeats went seven times. Analysis: Each occurrence describes a different movement. What to learn: Replacing went with hurried, wandered, climbed, crossed, returned, slipped and approached can increase diversity because the verbs encode different actions. The revision improves meaning and imagery at the same time. This is a good lexical-diversity intervention.
The false upgrade
Situation: A student replaces use with utilise everywhere. Analysis: The student assumes rarer vocabulary is more sophisticated. What to learn: Some replacements are grammatically possible but stylistically unnecessary; others sound bureaucratic. Lexical sophistication is not a contest for longer words. The ordinary verb may be clearer. Variety created by needless formalisation should not be rewarded.
The narrow evidence verbs
Situation: An analytical essay uses shows in every paragraph. Analysis: The writer needs a small functional network, not dozens of synonyms. What to learn: Teach indicates for pointing toward a pattern, suggests for cautious inference, demonstrates for strong support, illustrates for an example and reveals for information becoming visible. Diversity grows from epistemic precision.
The short-sample illusion
Situation: Two students each speak for twenty seconds. One produces twenty-five tokens and twenty-one types. Analysis: The resulting TTR looks exceptionally high. What to learn: The sample is too short for a stable conclusion. A few additional repetitions would change the ratio sharply. Collect longer or repeated samples before interpreting the apparent advantage.
The long-sample penalty
Situation: A student writes 1,000 words while another writes 200. Analysis: The longer essay has a much lower raw TTR. What to learn: The longer writer may still use far more unique words in absolute terms. The lower ratio partly reflects the mathematics of repetition over length. Compare fixed-length segments or use a more length-resistant measure.
The genre penalty
Situation: A laboratory report scores lower on diversity than a personal narrative by the same student. Analysis: The report repeats apparatus names, variables and procedures. What to learn: The difference may be a genre effect rather than a vocabulary decline. Technical genres reward stable terminology; narratives create more opportunities for varied action and description.
The topic-expertise boost
Situation: A football enthusiast speaks with high diversity about tactics but low diversity about architecture. Analysis: The person has richer conceptual networks in one domain. What to learn: Topic-specific knowledge changes the available lexical candidates. A general vocabulary judgement should sample more than one topic.
The bilingual distribution
Situation: A multilingual student gives a simple English explanation of a cultural practice but a rich explanation in another language. Analysis: English diversity alone looks low. What to learn: The conceptual knowledge is present; the lexical mapping differs across languages. Instruction can attach English forms and collocations to existing concepts rather than reteaching the concept from zero.
The prepared-speech effect
Situation: A presentation script has high lexical diversity, but the unscripted question period does not. Analysis: Prepared writing allows search and revision; spontaneous speaking requires fast retrieval. What to learn: The difference reveals an access gap. Productive training should include short, timed responses so the vocabulary becomes available without a script.
The word-bank effect
Situation: A student uses many target words when a word bank is visible. Analysis: Diversity appears to improve dramatically. What to learn: Remove the bank later. If the words disappear, the earlier score reflected cue availability rather than independent retrieval. Assessment must distinguish assisted from unassisted production.
The copied-phrase effect
Situation: A reading response contains diverse vocabulary lifted directly from the source text. Analysis: The response looks lexically rich. What to learn: Ask the student to explain the same ideas without looking. If the vocabulary disappears, textual uptake has not yet become owned productive vocabulary.
The pronoun solution
Situation: A report repeats the research team in every sentence. Analysis: Replacing the noun phrase with accurate pronouns or ellipsis can reduce surface repetition without inventing synonyms. What to learn: This is a discourse-level solution, not a vocabulary upgrade. Lexical diversity can sometimes improve through cohesion devices rather than new lexical knowledge.
The content-word audit
Situation: A student has normal overall TTR but repeats only three lexical verbs. Analysis: Function words and varied nouns hide the verb bottleneck. What to learn: A POS-specific audit reveals the problem. Targeting analytical verbs may improve writing more than adding more nouns.
The adjective flood
Situation: A creative paragraph uses many different adjectives. Analysis: The diversity score is high, but the prose feels overloaded. What to learn: Delete adjectives that repeat information already carried by the noun or scene. High diversity can coexist with verbosity. Good writing is selective.
The compact expert
Situation: An expert explanation uses few lexical types but each is precise and conceptually loaded. Analysis: A novice may use more varied everyday wording. What to learn: Lexical diversity alone can invert expertise. Domain knowledge often compresses explanation around stable technical terms.
The academic connector pattern
Situation: An essay repeats however and therefore. Analysis: The writer may lack discourse connectors, but variety is not the only issue. What to learn: First ask whether the logical relationships are genuinely contrast and consequence. Then expand with appropriate structures such as although, despite, as a result and consequently only where the logic fits.
The paraphrase challenge
Situation: Students rewrite one proposition three ways without changing meaning. Analysis: Successful versions use different vocabulary and syntax while preserving truth conditions. What to learn: This is stronger evidence of flexible lexical control than a high TTR because the task explicitly tests controlled alternatives.
The collocation trap
Situation: A learner replaces strong evidence with powerful evidence to avoid repetition. Analysis: The meaning is understandable but the collocation is less conventional. What to learn: Diversity should respect phrase naturalness. Collocation acts as a constraint on substitution.
The register trap
Situation: A student varies children with kids in a formal report. Analysis: The new word increases variation but shifts register. What to learn: Lexical appropriateness matters more than variety. Keep register stable unless the genre supports the change.
The repeated name
Situation: A biography repeats the person’s name frequently. Analysis: Pronouns could reduce some repetition, but excessive pronoun use may create ambiguous reference. What to learn: Lexical choice must be coordinated with discourse clarity. A simple metric cannot decide the best balance.
The rhetorical refrain
Situation: A speech repeats “We can do better” at the end of several sections. Analysis: The repeated phrase deliberately creates rhythm and emphasis. What to learn: A diversity metric marks repetition, but a rhetorical analysis recognises anaphora or refrain. Artistic purpose overrides numerical variety.
The young-child sample
Situation: A five-year-old tells a short story with high TTR. Analysis: The ratio is partly inflated by brevity. What to learn: Developmental interpretation should consider total vocabulary, grammatical growth, narrative structure and repeated samples rather than celebrate the ratio alone.
The research-method paragraph
Situation: A paper repeats participants, measure, trial and condition. Analysis: These nouns define the experimental structure. What to learn: Varying them could make the method harder to follow. Precision and reproducibility legitimately constrain diversity.
The coding tutorial
Situation: Technical instructions repeat exact interface labels and command names. Analysis: Consistency reduces lexical diversity. What to learn: Users need the same label they see on screen. Terminological variation would reduce usability and safety.
The legal clause
Situation: A contract repeats a defined term instead of substituting everyday synonyms. Analysis: The apparent repetition is a feature, not a flaw. What to learn: Legal interpretation depends on stable defined references. Lexical diversity is a poor quality target here.
The poetry case
Situation: A poem repeats one word obsessively. Analysis: The recurrence creates theme, sound and emotional pressure. What to learn: Quantitative lexical analysis can describe the repetition but cannot decide whether it is artistically successful.
The translation case
Situation: A translated text shows lower diversity than the source. Analysis: The target language may prefer repetition where the source tolerates substitution, or vice versa. What to learn: Cross-language comparison requires knowledge of linguistic and stylistic conventions, not only ratios.
The lemmatisation case
Situation: One tool counts run, runs, running, ran as different forms; another reduces them toward a lemma. Analysis: The diversity scores differ. What to learn: Neither is automatically wrong. They operationalise “different word” differently. Comparisons require consistent preprocessing.
The hyphenation case
Situation: A corpus counts decision-making as one token in one pipeline and two or more in another. Analysis: The result changes type and token totals. What to learn: Small preprocessing rules can matter, especially in short texts. Report them when precision matters.
The proper-name case
Situation: A narrative contains many place and character names. Analysis: Unique names inflate surface diversity. What to learn: If the research question concerns general lexical proficiency, analysts may need a principled policy for names rather than assuming each proper noun reflects vocabulary richness.
The spelling-error case
Situation: A misspelled word can be counted as a unique type by naïve software. Analysis: Errors artificially raise diversity. What to learn: Clean or annotate data according to a documented procedure before interpreting scores.
The inflection case
Situation: A morphologically rich language produces many surface forms from one lexical base. Analysis: Form-based diversity may rise because grammar creates variation. What to learn: Cross-language comparisons need methods sensitive to morphological structure.
The AI paraphrase case
Situation: A student asks an AI tool to “make this less repetitive.” Analysis: The output may display higher diversity immediately. What to learn: The change demonstrates tool capability, not necessarily student learning. Ask the student to explain why each replacement fits and reproduce the improvement later without the tool.
The rubric case
Situation: A marking rubric rewards “wide vocabulary.” Analysis: Students interpret this as “use difficult words.” What to learn: Rewrite the descriptor operationally: precise and appropriate word choice, sufficient range for the task, controlled repetition, natural collocations and suitable register. This protects students from gaming surface diversity.
The portfolio case
Situation: A student’s equal-length essays show gradually rising content-word diversity across six months. Analysis: At the same time, error rates fall and word choice becomes more precise. What to learn: Now the diversity trend is more persuasive because it aligns with other evidence and comparable sampling.
The intervention case
Situation: After vocabulary instruction, TTR rises but MTLD does not. Analysis: Different measures capture different aspects and react differently to sample structure. What to learn: Do not choose whichever metric supports the desired conclusion. Examine why the measures diverge and report the uncertainty.
The final judgement
Situation: Two texts have identical lexical-diversity scores. Analysis: One is precise and coherent; the other is full of awkward synonym substitutions. What to learn: The equal score proves the central point: lexical diversity is one property of language, not a general measure of writing quality.
39. Twenty measurement mistakes and how to repair them
Comparing unequal lengths
Mistake: A 100-word text and a 1,000-word text receive raw TTR scores and are ranked directly. Repair: Match sample length, use fixed-length windows or choose a measure designed to reduce length sensitivity. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Treating TTR as vocabulary size
Mistake: The analyst claims that a text with 80 types proves the writer knows 80 words. Repair: A sample reveals only the words produced there, not the full mental lexicon. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Treating diversity as sophistication
Mistake: A text with many unique common words is described as “advanced vocabulary.” Repair: Measure sophistication separately; variation and advancement are different constructs. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring accuracy
Mistake: Every unique word is counted positively even when several are misused. Repair: Combine diversity with semantic and grammatical accuracy. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring collocation
Mistake: Synonym substitutions raise diversity but create unnatural phrases. Repair: Audit phrase naturalness and register. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring genre
Mistake: Narratives and scientific reports are compared as if they offer identical lexical opportunities. Repair: Compare within genre or model genre effects. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring topic familiarity
Mistake: Students write about one familiar and one unfamiliar domain. Repair: Use multiple topics or control prior knowledge. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring revision conditions
Mistake: Untimed edited writing is compared with first-draft timed writing. Repair: Document production conditions and compare like with like. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring tool assistance
Mistake: One group uses dictionaries or AI and another does not. Repair: Assistance changes lexical search and should be treated as part of the task. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring word class
Mistake: An overall score hides severe verb repetition. Repair: Inspect noun, verb or content-word diversity when diagnostically relevant. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring proper nouns
Mistake: Unique names inflate diversity in narratives. Repair: Decide and report how names are handled. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring spelling variants
Mistake: Typos are counted as unique types. Repair: Normalise errors according to the analysis purpose. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring contractions
Mistake: Different tokenisers split or preserve forms differently. Repair: Use one consistent pipeline. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring lemmatisation
Mistake: Surface forms and lemmas are compared without distinction. Repair: Name the lexical unit being counted. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring morphology
Mistake: Languages with richer inflection are compared using raw word forms. Repair: Use language-appropriate methods. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Ignoring sample instability
Mistake: A twenty-word sample is given a precise numerical interpretation. Repair: Collect more language or repeat samples. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Choosing a metric after seeing results
Mistake: The analyst tries several indices and reports only the one with the desired pattern. Repair: Pre-specify the measure or transparently report multiple justified measures. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Assuming complex means valid
Mistake: A complicated index is treated as automatically superior. Repair: Validation depends on the task and sample, not mathematical complexity alone. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Turning feedback into score chasing
Mistake: Students are told to maximise lexical diversity. Repair: Teach precision, flexible retrieval and controlled repetition instead. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
Using diversity for authorship detection
Mistake: A threshold is used to accuse students of AI use. Repair: Lexical diversity is not a reliable authorship detector; use process evidence and transparent policy. The wider principle is to keep the construct, the sample and the decision aligned. A lexical-diversity number is useful only when the comparison is fair and the interpretation stays inside what the measure can support.
40. Ten extended teaching sequences
Four-week verb diversity programme
Plan: Week 1 diagnoses repeated general verbs in equal-length writing samples. Students sort high-frequency verbs by function: movement, reporting, causation, change, evaluation and cognition. Week 2 teaches small contrast sets such as show/indicate/suggest/demonstrate and cause/trigger/contribute/enable. Week 3 moves into timed retrieval and paragraph writing without word banks. Week 4 repeats the original writing condition and compares verb diversity, accuracy and meaning precision. Why it matters: The programme targets a specific bottleneck rather than telling learners to “use better vocabulary.” Success requires both increased range and correct deployment.
Noun precision programme
Plan: Students underline vague nouns such as thing, stuff, problem and idea. Each occurrence is classified: object, factor, issue, process, constraint, risk, claim, principle or consequence. Learners then build domain-specific noun networks from current school subjects and practise using them in explanations. Why it matters: Noun diversity becomes conceptual precision rather than decorative variation.
Adjective control programme
Plan: Students collect overused adjectives such as good, bad, big, small, nice and important. Instead of memorising synonyms, they organise alternatives by criterion: effectiveness, size, severity, value, probability, reliability, emotion and appearance. Each new adjective is tied to a natural noun phrase. Why it matters: The method prevents thesaurus errors by teaching semantic categories and collocation together.
Reporting-verb programme
Plan: Use short source extracts and ask what the author is doing: stating, claiming, arguing, observing, suggesting, demonstrating, warning or acknowledging. Students then write evidence sentences that preserve the correct strength of the source. Why it matters: This improves lexical range in academic writing while training evidence judgement.
Narrative movement programme
Plan: Collect common movement verbs and sort them by speed, direction, effort, stealth and emotional implication. Students act out or visualise movements, then write scenes where the verb replaces an adverb-heavy phrase only when appropriate. Why it matters: Diversity improves because verbs carry richer event meaning.
Lexical-diversity measurement lesson
Plan: Students calculate TTR on two equal 100-word passages, then on one passage after expanding it to 300 words. They observe the falling ratio and discuss why length matters. The teacher introduces MATTR conceptually as a moving fixed-size comparison. Why it matters: Students become metric-literate and are less likely to treat one ratio as a universal quality score.
Cohesion versus variety lesson
Plan: Give students a paragraph where every repeated key noun has been replaced by a different synonym. Ask them to identify whether the references still point to the same concept. Then restore strategic repetition and discuss clarity. Why it matters: Learners discover why lexical variety can conflict with cohesion.
Speech retrieval programme
Plan: Students prepare semantic clusters but practise with shuffled scenario cues. Responses are limited to a few seconds, then expanded into a complete sentence. The same target words return after increasing intervals and in different topics. Why it matters: The programme converts receptive options into fast productive access.
Reading-to-output programme
Plan: During reading, students collect only high-utility words that recur or carry important distinctions. They record phrase patterns and contrasts. Later they complete an unrelated writing task where the words could be useful but are not required. Why it matters: Spontaneous accurate use is stronger evidence of transfer than forced insertion.
Portfolio measurement programme
Plan: Across a term, students produce several equal-length pieces in the same genre under similar conditions. The teacher tracks repeated content words, POS-specific variety, lexical errors and writing quality. Trends are discussed alongside the texts rather than as isolated scores. Why it matters: Longitudinal evidence connects quantitative patterns to observable development.
41. Field guide: 25 decisions about repetition and variation
Repeat the exact word
Choose this when it is a defined technical term, central concept or unambiguous referent. The reason is because terminological stability protects precision and cohesion. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use a pronoun
Choose this when the referent remains unmistakable. The reason is because grammar can reduce clumsy repetition without inventing lexical substitutes. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use a near-synonym
Choose this when the meaning genuinely changes in strength, tone or perspective. The reason is because variation should encode a distinction. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use a more specific noun
Choose this when a vague word such as thing or problem hides the category. The reason is because specificity adds information. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use a stronger verb
Choose this when a general verb such as do or get conceals the process. The reason is because verbs can compress explanation. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Keep the simple word
Choose this when it is already exact and natural. The reason is because rarity is not a quality criterion. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Delete the repeated phrase
Choose this when repetition reflects redundancy rather than necessary reference. The reason is because compression can improve style more than substitution. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Restructure the sentence
Choose this when repetition comes from identical syntax. The reason is because lexical variety is not always the best repair. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Add an example
Choose this when abstract vocabulary remains vague. The reason is because concrete evidence can clarify meaning without changing the key term. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Add a contrast
Choose this when two words seem interchangeable. The reason is because semantic boundaries become visible through opposition. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Check a corpus or trusted dictionary
Choose this when collocation is uncertain. The reason is because phrase naturalness cannot always be inferred from definitions. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Preserve rhetorical repetition
Choose this when rhythm or emphasis is deliberate. The reason is because stylistic effect can outweigh surface diversity. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Vary reporting verbs cautiously
Choose this when sources make different epistemic moves. The reason is because argue, claim, observe and demonstrate do not mean the same thing. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Vary movement verbs
Choose this when events differ physically. The reason is because action-specific verbs improve narrative precision. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Vary evaluative adjectives
Choose this when different criteria are being judged. The reason is because effective, reliable and ethical express different evaluations. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Do not vary numerical labels
Choose this when categories or variables have fixed names. The reason is because measurement requires stable terminology. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Do not vary interface labels
Choose this when writing instructions for software or equipment. The reason is because users must match text to visible controls. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Do not vary legal defined terms
Choose this when exact reference has contractual consequences. The reason is because lexical creativity can create ambiguity. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use domain terminology
Choose this when the concept is established and the audience can understand it. The reason is because technical vocabulary compresses shared knowledge. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Explain domain terminology
Choose this when the audience may not know it. The reason is because precision without accessibility can block comprehension. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Prefer paraphrase over synonym swapping
Choose this when you need a different expression but no single equivalent is safe. The reason is because syntax and phrase structure can vary while meaning stays stable. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Use lexical chunks
Choose this when fluent language relies on conventional multiword expressions. The reason is because diversity can occur across phrases, not just individual words. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Allow repetition in speech
Choose this when it helps listeners track the message in real time. The reason is because spoken comprehension benefits from redundancy. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Reduce repetition in revision
Choose this when the writer has time to choose more precise alternatives. The reason is because editing can expose avoidable lexical narrowness. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
Assess later without support
Choose this when you want evidence of vocabulary learning. The reason is because assisted diversity may not survive independent production. The decision should be made from meaning, audience and genre first; a lexical-diversity score comes afterward as a description of the resulting language, not as the rule that generated it.
42. Glossary of 40 lexical-diversity and vocabulary-measurement terms
type
Definition: a distinct word form counted once within a sample. Why it matters: The definition depends on tokenisation and whether forms are lemmatised. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
token
Definition: one occurrence of a word form. Why it matters: Repeated occurrences each count as separate tokens. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
TTR
Definition: types divided by tokens. Why it matters: Easy to calculate but strongly affected by sample length. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
MATTR
Definition: average TTR across moving fixed-size windows. Why it matters: Reduces some length sensitivity; window size must be reported. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
MTLD
Definition: a measure based on how long diversity is sustained before crossing a threshold. Why it matters: Common in language research; interpretation still depends on sample quality. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
HD-D
Definition: a probability-based lexical-diversity measure derived from sampling ideas. Why it matters: One of several alternatives to raw TTR. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
VocD
Definition: a model-based measure designed to estimate lexical diversity across samples. Why it matters: Requires computational procedures rather than simple hand calculation. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
lemma
Definition: a base dictionary form grouping some inflected forms. Why it matters: Lemma-based diversity differs from surface-form diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
word family
Definition: a broader morphological grouping around a base or root. Why it matters: Useful in vocabulary-size research but not identical to a lemma. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
lexical sophistication
Definition: the advancement, rarity or contextual quality of lexical choices, depending on the framework. Why it matters: Different from diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
lexical density
Definition: the concentration of content words relative to function words. Why it matters: Different from diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
vocabulary breadth
Definition: the range or number of words known. Why it matters: A learner-level construct, not the same as sample diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
vocabulary depth
Definition: the quality and richness of knowledge about words. Why it matters: Includes dimensions such as meaning, morphology and collocation. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
lexical fluency
Definition: the speed and ease with which lexical knowledge is accessed and used. Why it matters: A person may know a word deeply but retrieve it slowly. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
content word
Definition: a noun, lexical verb, adjective or many adverbs carrying substantial lexical meaning. Why it matters: Definitions vary by analysis. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
function word
Definition: a grammatical word such as an article, auxiliary, pronoun or preposition. Why it matters: These repeat heavily and influence all-word diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
tokenisation
Definition: the procedure used to split text into countable units. Why it matters: Different tokenisers can produce different totals. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
lemmatisation
Definition: the procedure of reducing inflected forms toward a lemma. Why it matters: Useful when the research question concerns lexical choice beyond inflection. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
corpus
Definition: a structured collection of language samples. Why it matters: Corpora allow large-scale comparison of lexical patterns. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
window size
Definition: the number of tokens used in each moving MATTR calculation. Why it matters: A parameter that changes the resulting value. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
genre
Definition: a communicative text type such as narrative, report or argument. Why it matters: Genres create different lexical opportunities. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
register
Definition: language variation according to situation, audience and purpose. Why it matters: Lexical variety must remain appropriate to register. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
collocation
Definition: a conventional word partnership. Why it matters: Collocation constrains which synonyms can safely substitute. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
semantic precision
Definition: how exactly a word matches the intended meaning. Why it matters: A central quality dimension not measured by simple diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
cohesion
Definition: the linguistic links that help a text hang together. Why it matters: Strategic repetition can improve cohesion. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
reference
Definition: the relation between expressions and the entities or ideas they identify. Why it matters: Stable lexical reference can require repeated terms. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
repetition
Definition: recurrence of the same word or phrase. Why it matters: Can signal weakness, necessity, rhythm, cohesion or topic concentration. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
variation
Definition: use of different lexical forms across a sample. Why it matters: Variation is not automatically beneficial. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
sample length
Definition: the number of tokens or words in the analysed sample. Why it matters: One of the strongest influences on raw TTR. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
longitudinal analysis
Definition: comparison of language from the same learner across time. Why it matters: More persuasive when tasks and lengths are comparable. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
cross-sectional analysis
Definition: comparison among learners or groups at one time. Why it matters: Requires careful control of task and sample differences. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
productive vocabulary
Definition: words a learner can retrieve and use. Why it matters: A major source of observed lexical diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
receptive vocabulary
Definition: words a learner can understand when encountered. Why it matters: Can be much larger than the vocabulary visible in output. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
semantic network
Definition: connections among related concepts and words. Why it matters: Richer networks can create more flexible lexical choice. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
retrieval
Definition: accessing a word from memory when needed. Why it matters: Slow retrieval can narrow spoken diversity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
transfer
Definition: successful use of learning in a new task or context. Why it matters: Independent lexical variation can be evidence of transfer. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
metric validity
Definition: the degree to which an index supports the interpretation being made. Why it matters: A reliable number can still be invalid for the intended claim. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
reliability
Definition: the consistency of a measurement under comparable conditions. Why it matters: Consistency alone does not prove validity. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
construct
Definition: the underlying property a measure is intended to represent. Why it matters: Lexical diversity is one construct within broader lexical proficiency. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
proxy
Definition: an observable measure used as indirect evidence for something less directly observable. Why it matters: Diversity is sometimes used as a proxy for lexical range but must not be overinterpreted. In practice, always interpret the term inside the measurement framework being used rather than assuming every study uses identical operational definitions.
43. Twenty advanced practice prompts
Practice 1
Describe a crowded train station without using the words busy, people or went. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 2
Explain why a scientific result can be reliable without being valid. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 3
Compare two solutions using the verbs indicate, suggest and demonstrate accurately. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 4
Rewrite a paragraph that repeats good and bad by naming the actual evaluation criteria. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 5
Explain a technical process while deliberately repeating the key term for clarity. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 6
Write a 150-word narrative and then revise only vague repeated verbs. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 7
Give a one-minute explanation of climate adaptation, then repeat it using the same concepts but different phrasing. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 8
Paraphrase an argument three ways while preserving the strength of the original claim. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 9
Write a formal email and a casual message conveying the same request; compare vocabulary diversity and register. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 10
Describe a graph without using show more than once, but keep every replacement semantically accurate. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 11
Explain a mathematical procedure using precise operation verbs. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 12
Write a science paragraph where lower lexical diversity is desirable because terminology must stay stable. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 13
Create two 100-word texts with different TTR values but similar quality. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 14
Create two 100-word texts with similar TTR values but very different quality. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 15
Take a paragraph with forced synonyms and restore cohesion through strategic repetition. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 16
Build a verb network for cause: cause, trigger, contribute to, enable, lead to, accelerate. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 17
Build an evaluation network for important: central, relevant, necessary, influential, consequential, critical. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 18
Analyse a speech refrain and explain why repeated words can strengthen rhetoric. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 19
Compare spontaneous speech and revised writing on the same topic. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
Practice 20
Audit an AI-generated paragraph for unnecessary lexical variety. After completing the task, annotate every repeated content word and every deliberate substitution. For each change, state whether it improved semantic precision, collocation, register, cohesion or only surface variety. Then identify one word that should remain repeated because changing it would reduce clarity. This reflection is essential: the purpose is to train controlled lexical choice, not to maximise the number of unique types.
