THE MASTERY CLUB · VOCABULARY SERIES · LEXICAL PROFILING · VOCABULARY PROFILING · FREQUENCY BANDS · K-LEVELS · TEXT DEMAND
Vocabulary | Lexical Profiling — How to Measure the Vocabulary Demands of a Text
Lexical profiling, also called vocabulary profiling, is the process of analysing which kinds of words a text contains and how those words are distributed across frequency bands, vocabulary lists or other reference categories. A lexical profile can show how much of a text comes from the most frequent 1,000 word families, the second 1,000, academic vocabulary, mid-frequency bands, low-frequency vocabulary and off-list items. Search questions such as “What is a lexical profile?”, “How do I measure vocabulary difficulty in a text?”, “What are K1 and K2 words?”, “What is the Lexical Frequency Profile?”, “How do I use a vocabulary profiler?” and “How can I compare the vocabulary demands of two texts?” all point toward the same analytical job: make the hidden vocabulary composition of a text visible.
Lexical profiling is not the same as lexical coverage. A profile describes the text: which frequency levels, lists or lexical categories its words belong to. Coverage describes the relationship between a particular reader and that text: what proportion of the running words that reader knows. A text can contain 85% high-frequency vocabulary and still be easy for an expert who knows the specialist terms, or difficult for a learner who has gaps inside the high-frequency bands. Profiling is therefore a map of lexical demand, not a direct comprehension score.
This article is the canonical Lexical Profiling specialist inside Vocabulary — The Mastery Club apex. It is a completely new article and does not rewrite any existing eduKate page. It sits beside Vocabulary | Lexical Coverage, Vocabulary | Vocabulary Size, What Is Vocabulary | Lexical Diversity, What Is Vocabulary | Vocabulary Breadth and Vocabulary Depth and Vocabulary | Lexical Access. Its job is specific: explain frequency profiling, K-levels, word-family and lemma choices, academic and off-list categories, profiler tools, interpretation limits and practical uses for teachers, students, writers and text designers.
The 60-Second Answer: What Is Lexical Profiling?
A lexical profiler takes a text and compares its words against one or more reference lists. The output shows how many tokens, types, lemmas or word families fall into each category. In a classic profile, a text may be divided into the first 1,000 most frequent word families, the second 1,000, an academic list, and words not found on those lists. In newer systems, a profiler can report BNC/COCA frequency bands from K1 through K25, NGSL-based levels, academic spoken vocabulary, school vocabulary lists or other frameworks.
The classic Lexical Frequency Profile was proposed by Batia Laufer and Paul Nation in 1995 as a way to measure lexical richness in second-language writing through the proportion of high-frequency, academic and less frequent vocabulary. That idea became the foundation for many later web-based vocabulary profilers. Current Lextutor tools still implement the classic system while also offering BNC/COCA, NGSL and other list frameworks.
The key principle is simple: profiling converts a text from a stream of words into a lexical distribution. That distribution can help compare texts, tune reading materials, inspect learner writing and identify vocabulary bands that deserve attention. It cannot by itself tell you whether a student understands the text, whether the words are used accurately, or whether one rare word is more important than another.
Part I — The Four Questions a Lexical Profile Can Answer
1. How much of this text uses very high-frequency vocabulary?
The first 1,000 and second 1,000 frequency bands often account for a large proportion of ordinary English text. A profile makes that proportion explicit. This is useful when comparing graded readers, textbooks, news articles or learner writing.
2. How much academic or specialist vocabulary appears?
Depending on the list framework, a profiler may identify words from the Academic Word List, New Academic Word List, Academic Spoken Word List, Middle School Vocabulary List or a custom technical list. This reveals lexical layers that general frequency bands alone can hide.
3. How much vocabulary falls beyond the common bands?
Low-frequency and off-list items can signal technical terminology, proper nouns, spelling errors, names, new coinages or genuinely unusual vocabulary. The profiler tells you where to inspect, but human judgment must determine what those items mean for the text.
4. How does this text compare with another text?
If two texts are profiled with the same list framework and counting unit, their distributions can be compared. One may rely more heavily on K1–K2 vocabulary; another may contain more K5+ or academic items. This can support text sequencing, material selection and curriculum design.
Part II — The Classic Lexical Frequency Profile
5. K1: the first 1,000
K1 typically refers to the first 1,000 most frequent word families or lexical units in a particular frequency framework. These words usually provide a large share of running-text coverage because they occur so often. Exactly which items belong to K1 depends on the reference list and counting unit.
6. K2: the second 1,000
K2 contains the next frequency band. Together, K1 and K2 form the high-frequency core in many classic vocabulary models. In older Classic VocabProfile systems, K1 and K2 are paired with academic vocabulary and off-list items.
7. Academic vocabulary
The Classic profile historically used an academic list to identify words that occur across academic texts but are not fully captured by the most basic frequency bands. Modern profilers may offer the Academic Word List, Middle School Vocabulary List, Academic Spoken Word List, NAWL and other alternatives depending on age and purpose.
8. Off-list vocabulary
Off-list does not mean “bad,” “too difficult” or “unknown.” It means the item was not matched to the selected lists. Off-list items may be proper nouns, technical terms, spelling mistakes, compounds, abbreviations, new words, foreign-language items or vocabulary outside the list’s coverage. This category always requires inspection.
Part III — Modern K-Level Profiling
Modern profiling often extends beyond four broad categories. BNC/COCA-based tools can divide vocabulary into successive 1,000-family bands: K1, K2, K3 and onward. Lextutor’s current VP-Compleat interface offers BNC/COCA 1–25k families, lemma-based options, NGSL and other frameworks, allowing much finer-grained profiles than the original K1/K2/academic/off-list model.
9. Why more bands can help
A single “off-list” category collapses many different lexical levels. K3 vocabulary is not equivalent to K20 vocabulary. Fine-grained profiling can distinguish moderately frequent vocabulary from very rare vocabulary, which is useful for graded materials, corpus research and advanced learner writing.
10. Why more bands can also mislead
Fine-grained numbers create an illusion of precision if the underlying assumptions are ignored. A word’s band depends on corpus, variety, tokenisation, unit of counting and list construction. K7 in one framework is not automatically the same as K7 in another. Compare profiles only within the same system unless you understand the conversion problem.
Part IV — The Counting Unit Problem
11. Tokens
A token is one occurrence of a word in running text. If the appears fifty times, it contributes fifty tokens. Token percentages tell you what proportion of the reading stream belongs to each band.
12. Types
A type is a distinct written form. Repeating analyse ten times contributes one type but ten tokens. Type counts reveal lexical variety rather than running-text burden.
13. Lemmas
A lemma groups inflectionally related forms under a headword, depending on the framework. For example, walk, walks, walked, walking may be grouped under a lemma. Lemma-based profiling is often appropriate for productive vocabulary and modern corpus work.
14. Word families
A word family can group a headword with related inflected and derived forms under defined morphological rules. help, helpful, helpless, helpfully may belong to one family under a particular family system. Word families have been widely used for receptive vocabulary because readers can sometimes infer transparent derived forms from known roots and affixes.
15. Why the unit changes the result
A text profiled by word families can appear lexically easier than the same text profiled by lemmas because several derived forms collapse into one known family. A learner may know analyse yet not fully understand analytically. The counting unit therefore encodes an assumption about what counts as related lexical knowledge.
Part V — Corpus Choice Changes the Profile
Frequency is not a property floating inside a word. It is measured in a corpus. A corpus dominated by news, fiction, conversation, academic writing or one regional variety will produce different counts. A profiler based on BNC/COCA combines large British and American sources; other profilers use different corpora and lists. Recent 2026 research on Singapore English frequency and contextual diversity is another reminder that regional varieties can show systematic frequency differences.
16. General corpora
General corpora aim to represent broad language use across multiple genres. They are useful for general frequency bands and broad text comparison.
17. Academic corpora
Academic corpora reveal vocabulary common in research, textbooks or lectures. A word rare in conversation may be routine in academic prose.
18. Domain corpora
A medical, legal or engineering corpus changes what counts as frequent. Specialist learners should not judge domain vocabulary solely through general-English frequency.
19. Regional corpora
English varies geographically. Frequency, spelling, collocation and sense distribution can differ across varieties. A profile is always tied to the data behind its list.
Part VI — How a Vocabulary Profiler Actually Works
20. Step 1 — Tokenise the text
The profiler first separates running text into units that can be matched against a reference list. Seemingly trivial choices matter. Does don’t count as one token or two? What happens to hyphenated compounds? Are numbers discarded? Are apostrophes retained? Different tools make different tokenisation decisions, and those decisions can change the profile.
21. Step 2 — Normalise or parse forms
Some systems lower-case words, identify parts of speech, reduce inflected forms to lemmas, or map words into families. Others work directly with surface forms. A parsed profiler may distinguish noun record from verb record; a simpler system may not. The richer the preprocessing, the more assumptions enter the analysis.
22. Step 3 — Match against a reference list
Each lexical unit is compared with one or more lists. If the item belongs to the first 1,000 band, it is assigned K1; if to the next band, K2; if it belongs to an academic list, it may be assigned there; if nothing matches, it becomes off-list. Modern interfaces can apply different frameworks to the same text so users can compare how the analytical lens changes the result.
23. Step 4 — Count tokens, types, families or lemmas
The system records how many items fall into each category. Token counts describe the running text. Type counts describe distinct forms. Family or lemma counts describe grouped lexical units. A good profiler reports enough information for users to see what unit is being counted.
24. Step 5 — Convert counts to percentages
Percentages make texts of different lengths comparable. If 800 of 1,000 running tokens are K1, the K1 token percentage is 80%. If another text contains 74% K1 and more K4–K8 vocabulary, the second text may impose greater lexical demand under that framework.
25. Step 6 — Inspect the actual words
Never stop at the percentages. Open the word lists. A 7% off-list result could be mostly character names in a novel, chemical formulas in a science text, spelling errors in learner writing or genuine low-frequency vocabulary. The number tells you where to look; the word list tells you what happened.
Part VII — Current Vocabulary Profiler Frameworks
Current Lextutor VocabProfile tools demonstrate how far profiling has moved beyond the original four-bin model. The 2026 VP-Compleat interface offers multiple list frames, including BNC/COCA 1–25k families, Classic GSL/AWL, NGSL-based systems and other specialised lists. It also distinguishes families from lemmas and supports different grain sizes. This flexibility is powerful because different questions need different lexical lenses.
26. Classic GSL/AWL
The classic profile groups words into the first 1,000 general-service families, the second 1,000, an academic list and off-list vocabulary. It is historically important and remains useful for teaching and research where comparability with older work matters.
27. BNC/COCA 1–25k
BNC/COCA-based profiling extends the frequency scale to twenty-five 1,000-family bands. This helps distinguish K3 from K10 or K20 rather than throwing all less-common vocabulary into one off-list bucket. It is useful for graded reading research, advanced writing analysis and fine-grained text comparison.
28. NGSL-based profiling
The New General Service List and related lists provide another modern frequency framework. NGSL-based profiles may suit teaching contexts that prefer lemma-based modern general-service vocabulary rather than older General Service List families.
29. Academic spoken vocabulary
Academic language is not only written. Academic Spoken Word List frameworks can profile lectures, discussions and spoken academic texts, recognising that spoken academic English has its own distributional profile.
30. School vocabulary lists
Younger learners need frameworks calibrated to school texts rather than university corpora. Some profilers allow middle-school vocabulary lists or child-focused profiles. Age and curriculum change what counts as useful frequency information.
31. Custom technical lists
A teacher or researcher can sometimes supply a technical list so domain vocabulary is separated from generic off-list words. This is especially valuable in medicine, science, engineering and law where specialist terms would otherwise look simply “rare.”
Part VIII — The Proper Noun Problem
Proper nouns can distort profiles because names are often absent from general frequency lists. A novel containing Elizabeth, Darcy, Pemberley can accumulate off-list tokens that do not represent genuine vocabulary difficulty for a reader. News texts can contain countries, politicians, organisations and brands. Scientific texts contain species names and abbreviations.
32. Why names inflate off-list percentages
Frequency lists are designed mainly around lexical vocabulary, not every possible name. A profile that treats all proper nouns as unknown low-frequency vocabulary can make a simple text look difficult.
33. When names really are difficult
Proper nouns are not automatically easy. A reader may need to know who a historical figure is, which institution an acronym represents, or what place a name refers to. Profiling and comprehension are different. Recategorising names may improve lexical statistics while leaving background-knowledge demands untouched.
34. Practical handling
Decide before analysis whether proper nouns should remain off-list, be recategorised, or be analysed separately. Use the same rule across texts you want to compare. Lextutor interfaces include options for handling proper nouns precisely because the choice materially affects results.
Part IX — The Compound Problem
35. Transparent compounds
A compound such as school bus may consist of two high-frequency words whose combined meaning is easy. A profiler that treats the expression as two known tokens may approximate the learner experience reasonably well.
36. Lexicalised compounds
Expressions such as greenhouse effect or credit card can have conventional meanings larger than the components. Single-word profiling can underestimate phrase-level difficulty.
37. Hyphenated forms
Forms such as evidence-based, long-term, user-generated may be tokenised differently across tools. One profiler may split the components; another may leave the whole form off-list. Text preparation and software settings matter.
38. Technical compounds
Scientific English creates dense compounds such as gene-expression profile or high-temperature superconductor. Component words may be individually frequent while the conceptual combination remains technical. Profiling should be combined with domain analysis.
Part X — The Spelling and Formatting Problem
39. Spelling errors
A learner who writes enviroment may generate an off-list item even though the intended word is common. If the goal is to measure lexical sophistication rather than spelling, errors should be normalised or analysed separately. If the goal includes lexical form accuracy, preserving errors may be informative.
40. Capitalisation
Sentence-initial words, all-caps headings and inconsistent capitalisation can affect proper-noun detection or list matching. Clean text preparation reduces artefacts.
41. Numbers and symbols
Dates, mathematical notation, percentages and formulas can create noise. A science text with many numerical tokens may look lexically unusual even though numbers are not the vocabulary demand of interest. Decide how such items should be handled.
42. HTML and markup
Profiling raw web pages can accidentally count navigation labels, metadata, code and accessibility text. Extract the intended reading passage before profiling. The lexical profile should represent what the learner is expected to read.
Part XI — What K1, K2 and K3+ Percentages Mean
43. High K1 percentage
A high K1 token percentage means much of the running text comes from the most frequent band under the selected framework. This often makes the text more lexically accessible, but it does not guarantee simple syntax, familiar concepts or high comprehension.
44. High K2 percentage
K2 vocabulary adds less frequent but still broadly common vocabulary. In learner writing, increased use beyond K1 can sometimes reflect growing lexical range, but only if words are used accurately and appropriately.
45. Higher K-bands
K3, K4 and higher bands represent progressively less frequent vocabulary under a banded framework. More high-band vocabulary can indicate greater lexical sophistication, specialist subject matter, literary style or unnecessary obscurity. The profile alone cannot tell which.
46. Off-list percentage
Off-list percentage is a diagnostic flag, not a quality score. Inspect the items. Proper nouns, technical terms, foreign words, errors and neologisms can all contribute.
Part XII — Cumulative Coverage Curves
A useful modern output reports cumulative percentage: how much of the text is covered by K1 alone, K1–K2 together, K1–K3 together and so on. The curve shows how quickly the text approaches near-total lexical coverage as less frequent bands are added.
47. Fast-rising curve
If the cumulative profile reaches a very high percentage by K2 or K3, the text relies heavily on common vocabulary. This may suit beginning or intermediate readers, though conceptual difficulty can remain.
48. Slow-rising curve
If substantial coverage is still missing at K8 or K10, the text contains many less-frequent items. This may signal specialist content, advanced literary vocabulary or a text poorly matched to general-language learners.
49. Comparing curves
Two texts can have the same average word length but very different cumulative profiles. Frequency-based profiling therefore reveals lexical information that surface readability measures can miss.
Part XIII — Lexical Profiling vs Related Measures
| Measure | Main question | Unit of interpretation |
|---|---|---|
| Lexical profiling | Where the text’s vocabulary falls across frequency/list categories. | Text composition by lexical bands. |
| Lexical coverage | How much of the text a particular reader knows. | Reader × text relationship. |
| Vocabulary size | How many words or families a learner knows under a test definition. | Learner lexical breadth. |
| Lexical diversity | How varied the vocabulary in a text or sample is. | Repetition vs variety. |
| Vocabulary depth | How richly each word is known. | Quality of lexical knowledge. |
| Lexical access | How efficiently a known word can be retrieved. | Retrieval speed/availability. |
| Readability | How difficult a text may be under a broader formula or model. | Often combines sentence and word features. |
| Lexical density | How much content/lexical vocabulary is packed into the text. | Content-word concentration. |
| Word frequency | How often an item appears in a reference corpus. | A property of corpus counts, not a whole-text profile. |
Part XIV — Ten Good Uses of Lexical Profiling
- Compare two reading passages under the same frequency framework.
- Sequence texts from lexically simpler to lexically more demanding.
- Identify which frequency bands dominate a textbook chapter.
- Find academic or specialist vocabulary worth pre-teaching.
- Inspect learner writing for movement beyond the highest-frequency bands.
- Compare drafts across time while controlling for genre and task.
- Audit graded readers or instructional materials.
- Build vocabulary lists from actual curriculum texts rather than intuition alone.
- Identify off-list items that require human inspection.
- Support research on lexical richness, text difficulty and vocabulary development.
Part XV — Ten Bad Uses of Lexical Profiling
- Treating a higher K-band percentage as automatically better writing.
- Assuming off-list words are always difficult or sophisticated.
- Comparing outputs from different list frameworks as if the bands were identical.
- Ignoring text length and genre.
- Ignoring proper nouns, compounds and spelling errors.
- Assuming profile percentages equal learner comprehension.
- Using word-family results as if every derived form is known automatically.
- Ranking students by one short writing sample without controlling task.
- Replacing teacher reading of the actual text with profiler statistics.
- Turning corpus frequency into a rigid prescription against precise low-frequency vocabulary.
Part XVI — Worked Lexical Profiles: What the Numbers Actually Tell You
Text A — simple narrative
Illustrative profile: K1 89%, K2 6%, K3+ 3%, off-list 2%. Interpretation: Most running words come from high-frequency bands. The small off-list share may consist of names. This profile suggests modest lexical demand, but narrative inference and syntax still matter. Caution: Inspect the off-list names before calling the text ‘easy’.
Text B — general science explanation
Illustrative profile: K1 78%, K2 9%, K3–K5 6%, academic list 4%, off-list 3%. Interpretation: The text relies less heavily on the first 1,000 and more on mid-frequency and academic vocabulary. It likely requires broader lexical knowledge than Text A. Caution: Check whether off-list items are technical science terms rather than random rare words.
Text C — advanced literary passage
Illustrative profile: K1 76%, K2 8%, K3–K8 10%, off-list 6%. Interpretation: A larger low-frequency tail may indicate literary vocabulary, archaic forms, proper nouns or unusual style. Caution: Read the actual words before inferring ‘sophistication’.
Text D — technical engineering note
Illustrative profile: K1 72%, K2 7%, K3–K6 5%, technical list 12%, off-list 4%. Interpretation: General lexical bands understate the local importance of domain vocabulary. The text may be easy for engineers and difficult for general learners. Caution: Use a technical list or domain corpus rather than treating all specialist terms as noise.
Text E — learner essay 1
Illustrative profile: K1 93%, K2 4%, academic 2%, off-list 1%. Interpretation: The writer relies heavily on very common vocabulary. This might reflect limited lexical range, but it could also reflect a simple task or deliberately plain style. Caution: Compare with another essay of the same genre and length before evaluating development.
Text F — learner essay 2
Illustrative profile: K1 83%, K2 8%, academic 5%, off-list 4%. Interpretation: The writer uses more vocabulary beyond K1. This may indicate richer productive range if the words are accurate and appropriate. Caution: Inspect errors and collocations; higher-band use is not automatically better.
Text G — news report
Illustrative profile: K1 84%, K2 7%, K3–K5 3%, off-list 6%. Interpretation: Off-list share may be inflated by names, organisations, places and abbreviations. Caution: Recategorise proper nouns before comparing with non-news texts.
Text H — textbook glossary page
Illustrative profile: K1 55%, K2 5%, technical list 35%, off-list 5%. Interpretation: A glossary is lexically dense by design and should not be compared directly with continuous prose. Caution: Genre controls are essential.
Text I — spoken lecture transcript
Illustrative profile: K1 91%, K2 5%, academic spoken list 3%, off-list 1%. Interpretation: Spoken academic language can use very high-frequency grammar and still convey complex ideas through syntax and discourse. Caution: Use spoken academic lists and do not equate high K1 with intellectual simplicity.
Text J — social-media explainer
Illustrative profile: K1 90%, K2 5%, K3+ 2%, off-list 3%. Interpretation: Lexically accessible language can still contain complex claims, misinformation or compressed concepts. Caution: Profiling measures vocabulary distribution, not truth or reasoning quality.
Part XVII — How to Profile a Text for Teaching
- Choose the exact passage students will read, not the entire website or textbook.
- Clean away navigation, references and irrelevant metadata.
- Decide how to handle proper nouns, numbers and abbreviations.
- Choose a list framework appropriate to age and purpose.
- Choose the counting unit: family, lemma or surface form.
- Run the profile and record token percentages.
- Inspect cumulative K-level coverage rather than one band alone.
- Open the actual K3+, academic and off-list word lists.
- Separate domain terms from accidental noise.
- Identify words central to the lesson or text meaning.
- Check whether learners already know those words.
- Select a small number for pre-teaching or deep instruction.
- Keep the original text visible while interpreting the profile.
- Repeat the process with later texts using the same framework.
- Use student performance to test whether the lexical predictions were useful.
50. Why profiling before reading can improve planning
Teachers often discover vocabulary problems only after students struggle. Profiling can reveal that a supposedly accessible passage contains an unexpected cluster of K5–K8 words or a large technical layer. This allows support to be designed before the reading lesson.
51. Why profiling should not replace teacher knowledge
A profiler cannot know which words students learned last month, which concepts they already know in another language, or which term the exam requires. Statistical demand and actual learner demand overlap imperfectly.
Part XVIII — How to Profile Learner Writing
52. Control the task
Lexical Frequency Profile research found that genre control matters. A narrative and an academic argument naturally invite different vocabulary. Comparing a student’s two texts without task control can make normal genre differences look like vocabulary development.
53. Control text length
Very short samples are unstable. One unusual word can change percentages sharply. Longer samples generally provide more reliable lexical distributions, though the appropriate minimum depends on the measure.
54. Correct or preserve spelling deliberately
If the purpose is to measure vocabulary range, spelling errors can be normalised so known words do not become false off-list items. If form accuracy is part of the construct, keep a separate error analysis. Do not let software defaults decide the research question.
55. Inspect appropriate use
A student may increase K5 vocabulary by inserting rare words incorrectly. The profile will count the words but not judge collocation, register or meaning. Human evaluation remains necessary.
56. Track movement across several samples
A single profile is a snapshot. Repeated samples under comparable tasks can show whether a learner increasingly uses vocabulary beyond the most frequent bands while maintaining accuracy.
57. Compare profile with vocabulary-size measures
The original LFP idea was valuable partly because it connected vocabulary size with vocabulary use. A learner may know many words receptively but use a much narrower productive range. Combining direct vocabulary tests with writing profiles can reveal that gap.
Part XIX — Lexical Profiling and Writing Quality
Lexical profile data can correlate with proficiency under controlled conditions, but it should never be converted into a rule such as “more rare words = better essay.” Strong writing uses the right word at the right time. Plain high-frequency vocabulary can express difficult reasoning beautifully. Rare vocabulary can make writing less precise if chosen for display.
58. Sophistication is not rarity alone
A precise common verb can be better than an obscure synonym. Cause, show, use and change remain legitimate words. Lexical sophistication includes appropriateness, phraseology, semantic precision and control.
59. Academic words are not decorative upgrades
Words from academic lists should appear because the meaning and genre call for them. A writer who inserts moreover, aforementioned, facilitate without control may produce less natural prose despite a superficially higher academic profile.
60. Genre creates different optimal profiles
A personal narrative, laboratory report, policy brief and children’s story should not have the same lexical-frequency distribution. Quality is always evaluated relative to communicative purpose.
Part XX — Lexical Profiling and Text Difficulty
61. Frequency is one component of difficulty
Less frequent words generally increase lexical demand because fewer readers know them. Yet difficulty also depends on sentence structure, conceptual density, discourse organisation, background knowledge and phrase-level meaning.
62. A common-word text can be conceptually difficult
A philosophy passage can use mostly common words while expressing abstract relationships that are hard to understand. A profiler may report high K1 coverage without capturing conceptual complexity.
63. A rare-word text can be locally easy
A dinosaur enthusiast may know many globally rare species names. An engineer may know rare technical terms automatically. General frequency does not equal individual familiarity.
64. Repetition can reduce effective difficulty
A technical term may be rare globally but repeated twenty times within a chapter. Once learned, later occurrences become easy. A raw frequency profile does not model learning across the text unless the analyst interprets recurrence.
65. Morphology can reduce difficulty
A lower-frequency derivative may be transparent to a learner who knows its root and affixes. Word-family profiling tries to reflect this possibility, but morphological transparency varies. Family membership does not guarantee comprehension.
Part XXI — Lexical Profiling and Graded Readers
Graded readers deliberately control vocabulary to match learner levels. Profiling can audit whether a supposedly beginner text really stays inside a defined vocabulary range. It can also reveal where a story introduces off-level words.
66. Controlled vocabulary
Publishers may restrict headwords or frequency bands at lower levels. Profiling can verify how closely the actual text follows those constraints.
67. Proper nouns and story-specific words
A graded story still needs names and plot vocabulary. These may be off-list without creating serious learner difficulty. Human review separates necessary story words from accidental lexical overload.
68. Repetition as pedagogy
A graded reader may intentionally repeat new vocabulary. Type counts and token counts together reveal whether the text introduces many new items or repeatedly reinforces a smaller set.
Part XXII — Lexical Profiling and Textbook Design
69. Unit-level profiling
Profile several chapters to see whether lexical demand rises abruptly. A curriculum can unintentionally jump from heavily K1–K2 language to dense K6+ and technical vocabulary without enough support.
70. Cross-subject comparisons
Science, history and mathematics textbooks can have different lexical distributions. Profiling helps curriculum teams see where students encounter high academic or technical loads across the school day.
71. Vocabulary recycling
Compare units to identify whether important Tier 2 or technical vocabulary reappears. A curriculum that introduces terms once and never reuses them creates weak opportunities for consolidation.
72. Glossary audits
Check whether glossary entries actually correspond to high-impact unfamiliar words in the chapters. Some glossaries overfocus on obvious technical terms while missing high-utility academic vocabulary.
Part XXIII — Lexical Profiling and AI-Generated Text
AI systems can produce passages at requested reading levels, but the claimed level should be verified rather than trusted. A prompt asking for “simple vocabulary” may still generate low-frequency words, technical terms or phraseological difficulty. Profiling provides one external check.
73. Profile before giving generated text to learners
Run the passage through the same vocabulary framework used for other materials. Inspect outliers. Edit words that create unintended demand while preserving terms essential to the concept.
74. Do not optimise blindly for K1
Forcing every word into K1 can flatten precision and remove necessary subject terminology. The goal is accessible language around important concepts, not lexical impoverishment.
75. Compare human and AI drafts
Profiling can reveal whether an AI rewrite actually reduced lexical demand or simply shortened sentences. Combine lexical profiles with human reading and comprehension checks.
Part XXIV — The Vocabulary Demand Stack
| Layer | Typical profile signal | Why it matters |
|---|---|---|
| Layer 1 — High-frequency core | K1–K2 or equivalent | Provides most running words in many general texts. |
| Layer 2 — Mid-frequency general vocabulary | K3–K9 roughly, framework-dependent | Often separates intermediate from advanced reading demands. |
| Layer 3 — Academic vocabulary | AWL/NAWL/ASWL or school-academic lists | Supports formal explanation and school/university texts. |
| Layer 4 — Domain vocabulary | Technical/custom list | Carries specialist concepts. |
| Layer 5 — Proper nouns and entities | Names, places, organisations | May be lexically easy but knowledge-demanding. |
| Layer 6 — Off-list/novel items | Unmatched forms | Needs inspection for errors, neologisms, foreign items or rare vocabulary. |
| Layer 7 — Multiword expressions | Chunks/collocations | Often invisible to single-word profilers but crucial for real comprehension. |
Part XXV — Profile Interpretation Laboratory: 30 Texts
Beginner graded reader
Illustrative profile: K1 94%, K2 4%, K3+ 1%, off-list 1%. What it suggests: Strong high-frequency concentration. Likely suitable for fluent reading if syntax and story knowledge are accessible. What it does not prove: Check whether off-list items are names rather than difficult vocabulary.
Intermediate graded reader
Illustrative profile: K1 86%, K2 8%, K3–K4 3%, off-list 3%. What it suggests: More lexical stretch while retaining a high-frequency core. What it does not prove: Useful for learners moving beyond the first 2,000 families.
Advanced graded reader
Illustrative profile: K1 78%, K2 9%, K3–K6 8%, off-list 5%. What it suggests: The lower-frequency tail becomes substantial. What it does not prove: May need more selective glossing and pre-teaching.
Primary science text
Illustrative profile: K1 82%, K2 7%, school-academic 5%, technical 4%, off-list 2%. What it suggests: General language remains accessible while a small technical layer carries the science concepts. What it does not prove: Teach the technical words with diagrams rather than replacing them.
Secondary history text
Illustrative profile: K1 76%, K2 9%, academic 8%, domain 5%, off-list 2%. What it suggests: A notable academic layer increases reading demand beyond topic names. What it does not prove: Target cross-curricular verbs and abstract nouns alongside historical terms.
Secondary mathematics explanation
Illustrative profile: K1 81%, K2 6%, academic 4%, technical 7%, off-list 2%. What it suggests: Technical terms and symbolic language carry much of the subject load. What it does not prove: Profile only the prose; equations require separate analysis.
General newspaper article
Illustrative profile: K1 84%, K2 7%, K3–K5 3%, proper nouns/off-list 6%. What it suggests: Names may inflate off-list output. What it does not prove: Recategorise entities before comparing with textbooks.
Opinion editorial
Illustrative profile: K1 78%, K2 9%, K3–K7 8%, academic 3%, off-list 2%. What it suggests: Argumentative prose uses more mid-frequency vocabulary and stance language. What it does not prove: High-band words may reflect topic and style rather than quality.
Children’s nonfiction
Illustrative profile: K1 88%, K2 5%, academic 2%, technical 3%, off-list 2%. What it suggests: Accessible frame with a small concept vocabulary. What it does not prove: Good example of simple language around precise subject terms.
Research abstract
Illustrative profile: K1 65%, K2 8%, K3–K10 12%, academic 9%, technical/off-list 6%. What it suggests: Dense academic and specialist vocabulary. What it does not prove: Not suitable for general frequency judgments without discipline-aware interpretation.
Research methods section
Illustrative profile: K1 70%, K2 8%, academic 10%, technical 8%, off-list 4%. What it suggests: Methods prose often contains recurrent formal vocabulary and technical procedures. What it does not prove: A technical custom list can improve interpretation.
Literature extract
Illustrative profile: K1 80%, K2 8%, K3–K10 7%, off-list 5%. What it suggests: Rare descriptive words, archaic forms or names may produce a long tail. What it does not prove: Literary effect cannot be reduced to frequency bands.
Legal notice
Illustrative profile: K1 72%, K2 7%, K3–K8 6%, legal list 12%, off-list 3%. What it suggests: Specialist terminology is locally high-value despite low general frequency. What it does not prove: Do not simplify away words carrying legal effect.
Medical patient leaflet
Illustrative profile: K1 83%, K2 7%, academic 2%, medical 6%, off-list 2%. What it suggests: General language dominates, but several technical items may be high stakes. What it does not prove: Profile supports plain-language revision while preserving necessary medical terms.
Technical manual
Illustrative profile: K1 74%, K2 6%, K3–K6 5%, technical 13%, off-list 2%. What it suggests: Specialist terminology dominates local comprehension. What it does not prove: General frequency bands alone underdescribe expertise requirements.
Student narrative essay
Illustrative profile: K1 92%, K2 5%, K3+ 2%, off-list 1%. What it suggests: Heavy reliance on high-frequency language. What it does not prove: Could still be excellent if precise and coherent; do not penalise plain diction automatically.
Student argumentative essay
Illustrative profile: K1 84%, K2 8%, academic 5%, K3+ 2%, off-list 1%. What it suggests: More academic and mid-frequency vocabulary appears naturally with the task. What it does not prove: Compare only with similar genres.
Student essay with forced rare words
Illustrative profile: K1 76%, K2 8%, K5+ 10%, off-list 6%. What it suggests: Profile looks ‘advanced’ numerically, but unusual words may be inaccurate or showy. What it does not prove: Inspect collocations and register before praising sophistication.
Adult workplace email
Illustrative profile: K1 91%, K2 5%, professional 2%, off-list 2%. What it suggests: High-frequency vocabulary is normal for efficient workplace communication. What it does not prove: Do not confuse accessibility with simplicity of thought.
Policy report
Illustrative profile: K1 73%, K2 8%, academic 10%, technical 6%, off-list 3%. What it suggests: Formal reporting language creates a substantial academic layer. What it does not prove: Useful for profiling general professional reading demands.
Lecture transcript
Illustrative profile: K1 90%, K2 5%, academic spoken 3%, off-list 2%. What it suggests: Speech relies heavily on frequent grammar and discourse words. What it does not prove: Intellectual complexity may come through structure rather than rare words.
Podcast transcript
Illustrative profile: K1 93%, K2 4%, K3+ 1%, off-list 2%. What it suggests: Conversational lexical profile. What it does not prove: Names and slang can dominate off-list items.
Museum label
Illustrative profile: K1 77%, K2 6%, academic 5%, technical/proper 10%, off-list 2%. What it suggests: Short texts can be dense because every word carries content. What it does not prove: Text length makes percentages volatile.
Exam comprehension passage
Illustrative profile: K1 82%, K2 8%, K3–K6 6%, off-list 4%. What it suggests: Moderate lexical challenge with some less-common vocabulary. What it does not prove: Question language should be profiled separately if commands are a concern.
Exam question set
Illustrative profile: K1 88%, K2 7%, academic command words 4%, off-list 1%. What it suggests: Few words are difficult, but command vocabulary may be decisive. What it does not prove: Low overall demand can hide high task importance.
Website landing page
Illustrative profile: K1 92%, K2 4%, marketing terms 2%, off-list 2%. What it suggests: Designed for broad accessibility. What it does not prove: Raw HTML may contaminate the profile if not cleaned.
Product documentation
Illustrative profile: K1 79%, K2 6%, technical 11%, off-list 4%. What it suggests: Local technical terms matter more than general rarity. What it does not prove: Build a custom list for recurring product terms.
Scientific news explainer
Illustrative profile: K1 85%, K2 6%, academic 4%, technical 3%, off-list 2%. What it suggests: Popular science often wraps technical concepts in high-frequency language. What it does not prove: Good candidate for comparing expert vs public-facing prose.
AI-generated ‘Grade 6’ passage
Illustrative profile: K1 80%, K2 8%, K5+ 7%, off-list 5%. What it suggests: The requested grade label does not guarantee lexical calibration. What it does not prove: Edit out unintended high-band vocabulary but retain necessary concept words.
Bilingual learner’s essay
Illustrative profile: K1 88%, K2 6%, K3+ 3%, off-list/errors 3%. What it suggests: Off-list may mix spelling errors, names and L1-influenced forms. What it does not prove: Separate form errors from lexical-range analysis.
Part XXVI — A Teacher’s Lexical Profiling Workflow
- Select the exact learner group and reading purpose.
- Choose a profiling framework appropriate to age and context.
- Clean the text while preserving the actual language students will see.
- Decide how to handle names, numbers, abbreviations and compounds.
- Run the profile and save the summary.
- Inspect K1, K2 and cumulative K-level percentages.
- Inspect academic and technical list matches.
- Open the off-list items and classify them manually.
- Separate unavoidable domain terms from accidental difficult wording.
- Check which less-frequent words are central to comprehension.
- Compare the profile with what students have already learned.
- Choose pre-teaching targets.
- Decide which words deserve quick glosses instead.
- Run the lesson.
- Observe where students actually struggle.
- Revise your interpretation of the profile using learner evidence.
- Profile a later text with the same settings for comparison.
Part XXVII — A Student’s Lexical Profiling Workflow
Students do not need to profile every reading passage. The tool is most useful when they want to understand why a text feels difficult or why their writing feels repetitive.
- Paste a clean sample into one profiler.
- Record the chosen framework so future comparisons use the same one.
- Look at the first two or three frequency bands.
- Inspect the words beyond those bands.
- Mark which words you genuinely know and which only look familiar.
- Separate names and technical vocabulary.
- Choose a small number of useful unknown words to learn.
- If profiling your own writing, identify repeated K1 words that could be made more precise—not simply rarer.
- Check any unusual high-band word for natural collocation and register.
- Repeat on another comparable sample after several weeks.
Part XXVIII — Lexical Profiling for Curriculum Design
76. Map the year
Profile representative texts from each term. If lexical demand jumps suddenly, the curriculum may need stronger vocabulary preparation or more gradual sequencing.
77. Find recurring Tier 2 words
A word appearing across history, science and English may deserve coordinated instruction. Profiling multiple subjects can reveal cross-curricular vocabulary invisible inside one department.
78. Find recurring Tier 3 clusters
Technical words that repeat across a unit are strong candidates for pre-teaching, retrieval and glossary support. A profile helps distinguish one-off terms from conceptual anchors.
79. Detect vocabulary deserts
A curriculum may over-rely on simplified texts and expose students to too little mid-frequency or academic vocabulary. Profiling can reveal whether learners are being prepared for the lexical demands of later grades.
80. Detect lexical cliffs
The opposite problem occurs when texts jump too quickly into low-frequency and technical language. A lexical cliff can make subject learning look like a comprehension problem when the underlying issue is vocabulary demand.
Part XXIX — Profiling for Simplification and Adaptation
81. Simplify around the concept
If a technical term is essential, keep it. Simplify surrounding grammar or unnecessary rare vocabulary. The goal is access to the concept, not elimination of every low-frequency word.
82. Replace decorative rarity before technical necessity
A science explainer may contain both photosynthesis and ubiquitous. The technical term may be conceptually essential; the decorative adjective may be replaceable. Profiling helps identify candidates, while human judgment decides.
83. Protect key collocations
Replacing one word can break a natural phrase. Simplification should preserve collocations and technical phraseology where they carry stable meaning.
84. Re-profile after editing
Do not assume the revised version became easier. Run the same framework again and inspect the changed distribution. Then read the text normally to ensure meaning survived.
Part XXX — Profiling for Vocabulary List Creation
A profiler can generate data-driven candidate lists from actual texts. This is better than relying entirely on intuition, but list creation still requires filtering.
85. Exclude low-value proper nouns
Names can dominate off-list output without deserving vocabulary teaching. Keep entities only when the knowledge matters.
86. Prioritise recurrence
A word appearing ten times across a unit usually deserves more attention than an equally rare word appearing once.
87. Prioritise cross-text range
A word recurring across several texts or subjects has greater instructional leverage. Range can matter as much as raw frequency.
88. Separate technical and general targets
Technical terms should be learned with domain concepts; broadly useful words should be taught for transfer. Mixing them into one undifferentiated list weakens instruction.
89. Add phraseology
Profiling produces single-word candidates. Expand each important target with collocations, phrase frames and grammatical patterns before teaching it.
Part XXXI — Frequency Bands Are Descriptive, Not Moral
K1 is not “good vocabulary” and K12 is not “bad vocabulary.” Frequency bands describe distribution. The right word may be rare because the idea is precise. The wrong word may be common because it is vague. Educational use begins only after description is combined with purpose.
90. High frequency supports accessibility
Frequent words are generally known by more learners and processed more quickly. This makes them valuable for clear explanation.
91. Low frequency supports precision when needed
Rare terms can be indispensable. Photosynthesis, injunction and asymptote are not defects simply because they are uncommon in general corpora.
92. The best text uses the vocabulary its job requires
A public information leaflet and a doctoral methods paper should not share the same lexical profile. Appropriate profiling is always genre- and audience-sensitive.
Part XXXII — The Mathematics of a Lexical Profile
93. Token percentage
If a 1,000-token text contains 820 K1 tokens, K1 token coverage is 82%. The formula is simple: K1 tokens divided by total counted tokens, multiplied by 100. Token percentages describe how much of the reading stream comes from a band.
94. Type percentage
If the same text contains 400 distinct word forms and 220 of those are K1 types, K1 type percentage is 55%. Type percentage is usually lower than token percentage because high-frequency words repeat heavily. The difference between token and type profiles reveals repetition.
95. Cumulative percentage
If K1 covers 82%, K2 adds 8% and K3 adds 4%, cumulative K3 coverage is 94%. Cumulative profiles answer a planning question: how many frequency bands must be known before most of this text becomes lexically accessible under the framework?
96. Off-list percentage
If 30 out of 1,000 tokens are unmatched, off-list token percentage is 3%. Before calling those words rare, inspect them. Five proper nouns repeated six times could produce the entire 3%.
97. Family count vs token count
A profile may report fifty K3 tokens but only twelve K3 word families. This means a small number of K3 families are repeating. For teaching, twelve recurring families can be a much more manageable learning load than fifty unique unfamiliar items.
98. Coverage gain from learning one family
If one technical family accounts for twenty tokens in a chapter, mastering it can raise effective known-token coverage substantially. Profiling helps identify high-return words by combining band information with recurrence.
Part XXXIII — Cumulative Coverage Worked Examples
Profile 1
K1 88%, K2 7%, K3 2%, K4+ 1%, off-list 2% K1–K2 cumulative = 95%. The text may be accessible to readers with strong high-frequency vocabulary, assuming off-list items are manageable.
Profile 2
K1 79%, K2 8%, K3 5%, K4 3%, K5+ 3%, off-list 2% K1–K2 cumulative = 87%; K1–K4 = 95%. The text demands a broader mid-frequency vocabulary.
Profile 3
K1 71%, K2 7%, K3–K5 8%, K6–K10 6%, technical 6%, off-list 2% Even K1–K5 reaches only 86%. Specialist and lower-frequency vocabulary are structurally important.
Profile 4
K1 92%, K2 4%, K3+ 1%, names 3% General lexical demand is low; apparent off-list difficulty is mostly entities.
Profile 5
K1 85%, K2 5%, academic 7%, off-list 3% Academic vocabulary is the main extra layer. A school or university learner may benefit from targeted academic instruction.
Profile 6
K1 80%, K2 5%, domain 13%, off-list 2% General frequency looks moderate, but domain vocabulary dominates the remaining load. Expertise matters more than general word rarity.
Profile 7
K1 83%, K2 9%, K3 4%, K4+ 2%, off-list 2% K1–K2 already reaches 92%; modest expansion into K3–K4 may substantially improve access.
Profile 8
K1 75%, K2 7%, K3–K9 15%, off-list 3% A long mid-frequency tail signals a text likely written for advanced readers or a specialised audience.
Profile 9
K1 90%, K2 6%, K3+ 2%, off-list 2% Lexically simple distribution, but this could still be difficult if syntax or concepts are dense.
Profile 10
K1 65%, K2 8%, academic 10%, technical 12%, off-list 5% A highly specialised profile. General learners will need strong scaffolding even if the prose is short.
Part XXXIV — Word Families: Advantages and Limits
99. Why families are useful for receptive reading
If learners know a base word and productive morphology, they may infer related forms. A family-based profiler reflects this by grouping derivatives. This can approximate receptive reading more realistically than treating every form as completely unrelated.
100. Why families can overestimate knowledge
Morphological relationships vary in transparency. Knowing nation does not guarantee effortless knowledge of nationalise or every derived form. Some family members change meaning substantially. A family is a useful counting convention, not proof of learner knowledge.
101. Family definitions differ
Word-family systems depend on rules about which affixes and derivations are included. Different family lists can therefore assign forms differently. State the framework when reporting results.
Part XXXV — Lemmas: Advantages and Limits
102. Lemmas fit productive analysis
Productive writing requires control of specific lexical forms, so lemma-based measures can avoid assuming knowledge of distant derivations. Modern corpus lists frequently use lemmas or form-and-lemma hybrids.
103. Lemmas create larger vocabulary counts
Because derivational relatives are separated rather than collapsed, lemma-based systems produce more units than broad family systems. Vocabulary-size or text-demand figures cannot be compared directly across units without adjustment.
104. Part-of-speech ambiguity
Some lemma systems distinguish noun and verb uses; others group them. Record as noun and verb may be one lemma or two depending on the framework. Parsed profiling can handle this more precisely.
Part XXXVI — Surface Forms: Advantages and Limits
Surface-form profiling is transparent: every written form is counted as it appears. It requires fewer linguistic assumptions. However, walk, walks, walked, walking become separate units, which may exaggerate lexical variety or demand for learners who control inflectional morphology.
Part XXXVII — Choosing the Right List Framework
| Framework | Best use | Main caution |
|---|---|---|
| Classic GSL/AWL | Historical comparability; simple K1/K2/academic/off-list view. | Older list base and coarse off-list category. |
| BNC/COCA 1–25k families | Fine-grained frequency bands; useful for receptive vocabulary research. | Family assumptions; corpus blend may not match every variety or genre. |
| BNC/COCA lemma variants | More form-specific and suitable for productive analysis. | Larger unit count and less morphological compression. |
| NGSL | Modern general-service orientation; compact high-frequency core. | Different list philosophy means outputs are not directly equivalent to BNC/COCA bands. |
| NAWL / academic lists | Highlights cross-academic vocabulary. | Academic list membership is not the same as learner difficulty. |
| ASWL | Better suited to spoken academic language. | Not designed as a complete general-vocabulary system by itself. |
| School vocabulary lists | Age-appropriate school-text focus. | May not generalise to adult or university materials. |
| Custom technical list | Separates domain terms from general off-list vocabulary. | Quality depends on how the custom list was built. |
Part XXXVIII — Corpus Frequency Is Contextual
105. Genre frequency
A word can be frequent in fiction and rare in academic prose, or common in conversation and rare in legal writing. General-frequency lists average across genres. For specialised tasks, genre-specific frequency may be more informative.
106. Regional frequency
Different English varieties show different lexical habits. Recent work measuring Singapore English demonstrates why local corpora can add value: regional sociocultural patterns influence word frequency and contextual diversity. The same principle applies worldwide.
107. Historical frequency
Language changes. A list built decades ago can underrepresent newer vocabulary and overrepresent words whose usage has declined. Modern lists update the reference data but sacrifice some comparability with older studies.
108. Register frequency
A word can be common in formal writing but rare in speech. A single overall frequency rank can hide this. Choose a corpus matching the communication mode when possible.
Part XXXIX — Frequency vs Contextual Diversity
Raw frequency counts how often a word occurs. Contextual diversity asks across how many different contexts, documents or discourse environments it appears. A word repeated hundreds of times in one narrow domain can have high frequency but low range. A broadly distributed word may be more useful for general learners even with fewer total occurrences.
109. Why contextual diversity matters for learning
Encountering a word across different contexts can support richer representations and transfer. A profiler based only on global frequency cannot tell whether a word’s occurrences are widely dispersed.
110. Why range matters for Tier 2 selection
Broadly distributed words are strong candidates for high-utility instruction. This connects profiling with the Tier 1, Tier 2 and Tier 3 Vocabulary owner.
Part XL — The Profiler Does Not Know Meaning
111. Polysemy
A profiler may assign all instances of bank to one frequency category even when one means a financial institution and another a river bank. Frequency is typically form-based unless the system performs sense disambiguation.
112. Technical senses of common words
Power, function, mean, current, stress may be high-frequency forms but technical concepts in science or mathematics. Profiling can make a specialised text look easier than it is because the forms are common.
113. Multiword meaning
Take into account can consist entirely of common words while behaving as a conventional phrase. Single-word profiles miss phrase-level difficulty. Pair profiling with phraseological analysis when needed.
Part XLI — Lexical Profiling Glossary
Lexical profile
A distribution showing how words in a text map to frequency bands, lists or lexical categories.
Vocabulary profile
A common synonym for lexical profile in educational and applied-linguistics contexts.
Lexical Frequency Profile
The Laufer–Nation measure using proportions of vocabulary at different frequency/list levels to describe lexical richness, especially in L2 writing.
VocabProfile
A family of computer tools that match text words to vocabulary lists and frequency bands.
K-level
A 1,000-word or 1,000-family frequency band under a specific framework, such as K1 or K5.
K1
The first 1,000 most frequent lexical units under the chosen list system.
K2
The second 1,000 most frequent lexical units.
Cumulative coverage
The percentage of text covered when successive frequency bands are combined.
Token
One occurrence of a word in running text.
Type
A distinct surface word form.
Lemma
A headword grouping inflectional variants under a defined system.
Word family
A broader morphological grouping that may include derivatives as well as inflections.
Off-list
An item not matched to the selected reference lists.
Reference corpus
The collection of texts from which word frequencies or lists are derived.
Frequency list
A ranked or banded list based on corpus occurrence.
General Service List
A historically influential list of high-frequency general English words.
Academic Word List
A list of academic vocabulary designed to identify words frequent across academic texts beyond basic general-service vocabulary.
NGSL
New General Service List, a modern general-frequency list used in some profiling systems.
NAWL
New Academic Word List, used alongside NGSL-style frameworks.
ASWL
Academic Spoken Word List, designed for spoken academic vocabulary.
BNC
British National Corpus.
COCA
Corpus of Contemporary American English.
BNC/COCA lists
Frequency lists derived from combined or coordinated BNC and COCA evidence, widely used in vocabulary research.
Contextual diversity
How widely a word occurs across different contexts or documents rather than its raw total frequency alone.
Range
The spread of a word across texts, genres or domains.
Technical list
A custom or published list of specialist vocabulary from a domain.
Proper noun handling
The decision about how names are counted or recategorised in a profile.
Tokenisation
The process of dividing text into countable units.
Lemmatisation
Mapping inflected forms to lemmas.
Parsing
Assigning grammatical categories or structures to forms, sometimes before profiling.
Normalisation
Cleaning or standardising text forms before analysis.
Grain size
How fine or coarse the frequency divisions are, such as 1,000 bands versus smaller groups.
Lexical demand
The vocabulary knowledge a text is likely to require, interpreted from more than profile percentages alone.
Lexical richness
A broad construct involving range, sophistication, diversity and other qualities of vocabulary use.
Lexical sophistication
Use of vocabulary beyond highly frequent words, interpreted together with appropriateness and control.
Lexical diversity
Variation in vocabulary rather than frequency-band position.
Lexical density
Concentration of content words in text.
Lexical coverage
Proportion of text known by a particular reader, not the same as profile distribution.
Readability
Broader estimate of text difficulty, often combining lexical and sentence features.
Off-list noise
Items unmatched for reasons unrelated to true rarity, such as names, errors or formatting.
Profile comparability
The requirement that texts be analysed with compatible frameworks, units and preparation rules before percentages are compared.
Part XLII — Frequently Asked Questions About Lexical Profiling
What is lexical profiling?
Lexical profiling is the analysis of a text’s vocabulary distribution against reference lists or frequency bands. It shows how much of the text comes from high-frequency, academic, mid-frequency, technical or off-list vocabulary.
What is vocabulary profiling?
Vocabulary profiling is commonly used as a synonym for lexical profiling. In educational settings it usually means matching words in a text against frequency or vocabulary lists to produce a profile.
What is a Lexical Frequency Profile?
The Lexical Frequency Profile, associated with Laufer and Nation’s 1995 work, describes lexical richness through proportions of words from different frequency/list categories. It was originally developed especially for analysing L2 writing.
What does K1 mean?
K1 usually means the first 1,000 most frequent word families or lexical units under a particular profiling framework. The exact members depend on the list system.
What does K2 mean?
K2 usually means the second 1,000 most frequent units. K1 and K2 together often represent a high-frequency core.
What does K3 mean?
K3 is the third 1,000 band under a banded frequency system. Higher K numbers generally contain less frequent vocabulary.
Does K1 mean easy vocabulary?
Not automatically. K1 words are frequent, but a common word can have an unfamiliar sense or occur inside a difficult phrase or concept.
Does K10 mean difficult vocabulary?
It often signals lower-frequency vocabulary, but difficulty depends on the learner and context. A specialist may know K10 domain vocabulary better than a K3 general word.
What is off-list vocabulary?
Words or forms that the selected profiler cannot match to its lists. Off-list items may be rare words, proper nouns, technical terms, spelling errors, abbreviations, compounds or foreign-language items.
Is off-list vocabulary bad?
No. It is simply unmatched under the chosen framework. Some off-list words are essential technical terms or names.
What is cumulative coverage?
The percentage of running text covered when successive frequency bands are combined, such as K1+K2+K3.
What is the difference between a lexical profile and lexical coverage?
A lexical profile describes the distribution of vocabulary in a text. Lexical coverage describes how much of that text a particular reader knows.
What is the difference between lexical profiling and vocabulary size?
Profiling analyses a text or writing sample. Vocabulary size estimates how many lexical units a learner knows.
What is the difference between lexical profiling and lexical diversity?
Profiling focuses on frequency/list categories. Lexical diversity focuses on how varied the vocabulary is and how much repetition occurs.
What is the difference between profiling and readability?
Readability is a broader attempt to estimate text difficulty, often using sentence length, word length or statistical models. Lexical profiling isolates the vocabulary distribution.
What is the difference between a word family and a lemma?
A lemma usually groups inflectional forms, while a word family can also group derived forms under defined morphological rules. Exact definitions vary by framework.
Why do word families matter?
They reflect the idea that readers can sometimes understand derived forms through known roots and morphology. They reduce the number of counted units compared with lemma-based systems.
Why do lemmas matter?
Lemmas are often better suited to productive analysis because they do not assume knowledge of a large set of derivational relatives.
Can I compare a family-based profile with a lemma-based profile?
Not directly. The counting units differ, so percentages and band membership can shift. Compare like with like unless you explicitly model the difference.
Why do proper nouns cause problems?
Names are often absent from frequency lists and therefore appear off-list even when they pose little lexical difficulty. They should be inspected or handled consistently.
Why do compounds cause problems?
Profilers may split or fail to recognise compounds, while the phrase may have a conventional meaning beyond its parts. Tokenisation rules change results.
Do spelling mistakes affect profiling?
Yes. Misspelled common words can appear off-list. Decide whether to correct them depending on whether the goal is lexical range or form accuracy.
Can lexical profiling measure reading difficulty?
It can measure one important component—vocabulary distribution—but not syntax, background knowledge, inference, cohesion or conceptual density.
Can lexical profiling predict comprehension?
It can contribute to predictions, especially when combined with learner vocabulary data, but profile percentages alone do not guarantee comprehension.
Can lexical profiling measure writing quality?
Not by itself. It can describe lexical distribution and sometimes relate to proficiency under controlled conditions, but accurate, appropriate and coherent word use still requires human judgment.
Can more low-frequency words improve an essay?
Only if they are precise, natural and appropriate. Rare words inserted for display can reduce writing quality.
Can a simple lexical profile describe an academic paper?
It can provide useful frequency information, but academic and technical lists are often needed to interpret specialist vocabulary properly.
What is the best profiler?
There is no universal best tool. Choose based on age, language variety, research question, counting unit and list framework. Current Lextutor tools offer several useful options.
What is VocabProfile?
VocabProfile refers to web-based vocabulary profiling tools derived from the Laufer–Nation tradition, including Lextutor implementations that match text against frequency and academic lists.
What is VP-Compleat?
VP-Compleat is a current Lextutor vocabulary profiler offering multiple frameworks such as BNC/COCA, Classic GSL/AWL and NGSL-based lists with different units and grain sizes.
What is the Classic VocabProfile?
The Classic system typically reports K1, K2, academic-list and off-list categories and is closely connected to the historical Lexical Frequency Profile tradition.
What are BNC/COCA bands?
They are frequency bands derived from British National Corpus and Corpus of Contemporary American English data under specific list-construction methods. BNC/COCA family lists can extend to K25.
What is NGSL?
The New General Service List is a modern high-frequency general-English list used in some profiling frameworks.
What is AWL?
The Academic Word List is a widely used list of academic word families developed to capture vocabulary common across academic texts beyond a basic general-service core.
What is NAWL?
The New Academic Word List is a later academic vocabulary list designed to work with newer general-service list frameworks.
What is ASWL?
The Academic Spoken Word List targets vocabulary in spoken academic discourse.
Can I use a profiler for Primary students?
Yes, but choose an age-appropriate framework and interpret names, school vocabulary and morphology carefully. Child-focused profilers or school lists may be more appropriate than university-oriented lists.
Can I use a profiler for Secondary students?
Yes. Profiling can help calibrate textbooks, comprehension passages and learner writing, especially when paired with school-academic or BNC/COCA bands.
Can I use a profiler for JC or university students?
Yes. Fine-grained frequency bands and academic lists are especially useful for advanced reading and writing analysis.
Can I profile spoken transcripts?
Yes, but spoken academic or conversational frequency frameworks may be more appropriate than written academic lists. Transcripts also contain fillers, contractions and discourse markers.
Can I profile AI-generated text?
Yes. Profiling is a useful way to verify whether a generated passage actually meets the intended vocabulary level, though human review is still necessary.
Can I profile a whole novel?
Yes with tools that support large input, but interpretation should consider names, repeated story vocabulary and genre. Some web interfaces have input limits.
Can I profile a textbook chapter?
Yes. Clean the text, decide how to handle diagrams and formulas, then inspect technical terms separately.
Can I profile a list of words?
Yes, but token percentages behave differently when every item appears once. A word list is not running prose, so interpret the output accordingly.
Can I use profiling to choose vocabulary to teach?
Yes. Profiling can identify less-frequent, academic and technical words, but teacher judgment should decide which are important, useful and unknown.
Can profiling find Tier 2 vocabulary?
It can generate candidates by showing broadly less-common or academic words, but Tier 2 depends on instructional utility and learner context, not frequency bands alone.
Can profiling find Tier 3 vocabulary?
A custom technical list or domain analysis can identify specialist terms, but off-list status alone is not enough to classify Tier 3.
What does a high K1 percentage mean?
The text relies heavily on very frequent vocabulary. This generally supports lexical accessibility but says little about conceptual or syntactic difficulty.
What does a low K1 percentage mean?
More of the text comes from less frequent vocabulary. This can increase lexical demand, though domain expertise may compensate.
What does a high off-list percentage mean?
It means many tokens were not matched. Inspect them before interpreting: they may be names, errors, technical terms or genuinely rare vocabulary.
How should I handle proper nouns?
Either keep them as off-list, recategorise them, or analyse them separately—but use the same policy when comparing texts.
How should I handle spelling errors in learner writing?
If measuring lexical range, consider correcting obvious errors before profiling and record errors separately. If form accuracy is part of the construct, preserve them for that analysis.
How should I handle hyphenated words?
Check how your profiler tokenises them. For comparisons, use the same preprocessing rules across texts.
How should I handle abbreviations?
Decide whether they are technical vocabulary, names, symbols or noise. A profiler may classify them off-list without understanding their role.
Why can two profilers give different results?
They may use different corpora, lists, tokenisation rules, units, proper-noun handling and frequency-band definitions.
Should I always use the latest list?
Not necessarily. Use a modern list for current teaching when appropriate, but use the historical framework if comparability with earlier research matters.
What is the biggest mistake in lexical profiling?
Treating the output as self-interpreting. Percentages are only meaningful when you know the list, unit, text preparation, genre and analytical purpose.
What is the second biggest mistake?
Treating lower-frequency vocabulary as automatically better or harder. Frequency is descriptive; appropriateness is contextual.
What is the most useful classroom question after profiling?
Which words in this profile actually block these learners from understanding or using this text?
What is the most useful writing question after profiling?
Are the less-frequent words accurate, natural and necessary, or are they decorative substitutions?
What is the main takeaway?
Lexical profiling makes vocabulary distribution visible. Its value comes from combining that map with learner knowledge, text purpose and human interpretation.
Part XLIII — Research Boundary
The Lexical Frequency Profile was introduced by Laufer and Nation in 1995 as a measure of lexical richness in L2 writing. Their study found that, with genre control, profile patterns corresponded with vocabulary size and could discriminate among proficiency levels. That historical result made frequency profiling influential in vocabulary research.
Modern profiling has expanded the original idea rather than replacing it. Current Lextutor tools provide Classic, BNC/COCA, NGSL and other list systems, and current interfaces distinguish families from lemmas and support fine-grained bands. These improvements make profiling more flexible while increasing the importance of reporting exactly which framework was used.
Frequency profiling remains one measure among many. Research on lexical richness also uses diversity, sophistication, density, contextual diversity and other indices. No single statistic fully captures word knowledge or writing quality. The responsible approach is construct-first: decide what you want to measure, then choose the profile that actually represents that construct.
- Laufer & Nation (1995) — Vocabulary Size and Use: Lexical Richness in L2 Written Production
- Lextutor — VocabProfilers home
- Lextutor — VP-Compleat current interface
- Lextutor — Classic English VocabProfile
- Behavior Research Methods (2026) — word frequency and contextual diversity measures for Singapore English
Part XLIV — What This Article Does Not Claim
It does not claim that K1 words are always easy. It does not claim that K10 words are always sophisticated. It does not claim that off-list words are unknown. It does not claim that one profiler is universally superior. It does not claim that word families perfectly represent learner knowledge. It does not claim that lexical profiles replace comprehension testing, teacher judgment or close reading.
It claims something narrower and more useful: a vocabulary profiler can reveal the lexical distribution of a text or writing sample, and that distribution can improve material selection, vocabulary planning and research when the framework and limitations are understood.
Part XLV — Lexical Profiling Workbook: 50 Diagnostic Decisions
1. A text has 8% off-list items, but six character names repeat throughout the story.
Interpretation: Recategorise or separate proper nouns before interpreting lexical difficulty.
2. A learner essay has 5% off-list items caused mostly by spelling errors.
Interpretation: Correct obvious spelling if the goal is lexical range; analyse form errors separately.
3. A science text has only 2% off-list vocabulary but students struggle badly.
Interpretation: Check technical senses of common words, syntax and background knowledge. Profiling may be underestimating conceptual demand.
4. A legal text has 12% technical-list vocabulary.
Interpretation: Do not simplify away legally necessary terms; scaffold definitions and preserve legal effect.
5. Two passages are profiled with different word-list frameworks.
Interpretation: Do not compare percentages directly. Re-profile both under one framework.
6. One text is 100 words and another 2,000 words.
Interpretation: Short-text percentages are more volatile. Interpret the 100-word profile cautiously.
7. A student increases K5+ use between two essays, but the second essay is a different genre.
Interpretation: Genre may explain the change. Compare like tasks before attributing development.
8. A graded reader claims beginner level but shows a long K6+ tail.
Interpretation: Inspect the K6+ items and decide whether they are names, repeated story terms or accidental lexical overload.
9. A learner knows a technical term repeated twenty times.
Interpretation: General frequency may classify it low, but effective local coverage is high for that learner.
10. An AI-generated passage has simple sentences but 9% K7+ vocabulary.
Interpretation: Sentence simplification did not guarantee lexical simplification. Edit the vocabulary layer.
11. A text has 95% K1–K2 cumulative coverage.
Interpretation: This may support accessibility for readers with strong high-frequency vocabulary, but it is not the same as 95% known coverage for a particular learner.
12. A research abstract has high academic-list coverage.
Interpretation: Expected for the genre. Do not treat it as automatically better writing than plain prose.
13. A student essay contains many academic-list words used incorrectly.
Interpretation: Profile counts cannot judge correctness. Inspect collocation and meaning.
14. A news article has 7% off-list items that are organisations and place names.
Interpretation: Entity density, not rare general vocabulary, is driving the off-list score.
15. A technical manual shows low K1 but high custom-list coverage.
Interpretation: The text is specialist rather than randomly rare. Audience expertise is central.
16. A passage has high K1 but contains several phrasal expressions unfamiliar to learners.
Interpretation: Single-word profiling misses phraseological demand. Add chunk analysis.
17. A learner’s writing uses only K1–K2 vocabulary but is precise and coherent.
Interpretation: Do not lower the writing score merely because the profile is common-word heavy.
18. A learner’s writing uses K10 words unnaturally.
Interpretation: Higher bands do not equal quality. Correctness and register matter more.
19. A novel contains archaic vocabulary.
Interpretation: Frequency bands can flag rarity but not historical or stylistic purpose.
20. A Primary text contains many transparent compounds.
Interpretation: Check tokenisation before interpreting off-list output.
21. A science chapter contains formulas and units.
Interpretation: Remove or analyse symbols separately if lexical demand is the research target.
22. A transcript contains many fillers such as um and you know.
Interpretation: Spoken profiling needs preprocessing choices and may require a spoken-language framework.
23. A student uses derived forms such as analytically after learning analyse.
Interpretation: Family-based profiling may group them; lemma-based profiling may distinguish them. Choose according to the construct.
24. Two tools disagree about the band of one word.
Interpretation: Check corpora, list versions and counting units. Different profiles can both be internally correct.
25. A vocabulary list is profiled as running text.
Interpretation: Remember that token percentages lose their normal meaning when every item appears once.
26. A short exam instruction contains only common words but students misunderstand it.
Interpretation: Task meaning, not lexical rarity, may be the problem. Profile is not a full comprehension model.
27. A textbook unit repeats the same ten K4 words.
Interpretation: This may be desirable recycling rather than excessive difficulty.
28. A unit contains fifty different K4 words once each.
Interpretation: The lexical learning burden is much larger even if the token percentage resembles the previous unit.
29. A passage contains function many times in mathematics.
Interpretation: The form may be high-frequency generally but its mathematical sense is technical. Form-based profiling can hide sense difficulty.
30. A learner’s profile improves after memorising rare synonyms.
Interpretation: Check whether those words are used naturally and whether reading/writing capability actually improved.
31. A teacher wants to compare two versions of a simplified text.
Interpretation: Profile both using identical settings and inspect which low-frequency items were removed or retained.
32. A curriculum team profiles science, history and English units.
Interpretation: Use the same framework to find cross-curricular academic words and separate subject-specific technical layers.
33. A student asks which K-band they should memorise next.
Interpretation: Do not prescribe bands mechanically. Combine vocabulary-size evidence, reading goals and actual text demands.
34. A profiler labels a brand name as off-list.
Interpretation: Decide whether the brand matters to understanding; rarity alone is irrelevant.
35. A text has a high percentage of K1 function words.
Interpretation: This can happen in both simple and complex prose. Grammatical density and conceptual structure still matter.
36. A poetic text uses common words metaphorically.
Interpretation: Frequency profiling may classify the vocabulary as easy while interpretation remains difficult.
37. A learner recognises K5 words in reading but never uses them in writing.
Interpretation: Profiling receptive text demand and productive writing are different jobs. Use vocabulary size/access measures as well.
38. A teacher profiles a read-aloud book above students’ decoding level.
Interpretation: Lexical profiling can still help plan oral vocabulary support because listening comprehension also depends on word knowledge.
39. A technical word appears off-list but is defined clearly in the passage.
Interpretation: The text may teach its own vocabulary. Difficulty depends on how well the definition and context work.
40. A text uses many low-frequency words that are cognates for the learners.
Interpretation: General frequency may overpredict difficulty for that language group.
41. A text uses common English words with culture-specific references.
Interpretation: General frequency may underpredict knowledge demands.
42. A student writes many proper nouns in a history essay.
Interpretation: Proper nouns can inflate off-list percentages without indicating lexical sophistication.
43. A learner profile is generated from 80 words of writing.
Interpretation: Treat it as exploratory only; short samples are unstable.
44. A school wants one profiling framework from Primary to university.
Interpretation: One framework may aid continuity, but age-appropriate lists can improve interpretability. Decide whether longitudinal comparability or local fit matters more.
45. A text shows 98% cumulative coverage by K5.
Interpretation: This means 98% of tokens fall within the first five bands, not that a learner knows 98% of the text.
46. A passage has 3% technical words but each is central to the main idea.
Interpretation: Small percentages can still represent major conceptual load.
47. A passage has 8% low-frequency descriptive vocabulary but the plot remains clear.
Interpretation: Lexical rarity may affect style more than core comprehension.
48. A profiler cannot recognise a new technology term.
Interpretation: Off-list status reflects list lag as much as lexical rarity.
49. A learner’s K1 percentage falls over a year while writing quality improves.
Interpretation: This may reflect broader vocabulary use, but verify accuracy, task comparability and genre.
50. A learner’s K1 percentage stays stable while vocabulary knowledge improves.
Interpretation: Growth can appear in depth, phraseology, precision and comprehension without dramatic profile change.
Part XLVI — A Material-Selection Comparison Template
When comparing two texts, record the following fields. The template forces analysts to keep methodological decisions visible.
- Text title and genre
- Intended learner age and proficiency
- Total token count
- Profiler and version
- Reference list framework
- Counting unit: family, lemma or surface form
- K1 token percentage
- K2 token percentage
- Cumulative K3/K5/K10 coverage
- Academic-list percentage
- Technical-list percentage
- Off-list percentage
- Number of proper nouns
- Number of unique technical terms
- Repeated high-impact words
- Phraseological difficulties not captured by single-word profiling
- Known background-knowledge demands
- Observed learner comprehension
- Final teaching decision
Part XLVII — A Learner-Writing Comparison Template
- Writing task and genre
- Prompt
- Time limit and support conditions
- Word count
- Spelling normalisation policy
- Profiler framework
- K1/K2 distribution
- Academic vocabulary share
- Higher-band distribution
- Off-list errors vs legitimate items
- Lexical diversity measure, if used
- Accuracy of lower-frequency words
- Collocational control
- Register appropriateness
- Comparison with a prior same-genre sample
- Interpretation of change
Part XLVIII — A Text Simplification Audit
- Profile the original text.
- Mark words above the intended frequency range.
- Separate necessary technical terms from replaceable rare vocabulary.
- Identify common words used in specialist senses.
- Identify multiword expressions likely to be difficult.
- Rewrite only where access improves without losing conceptual precision.
- Profile the revision using identical settings.
- Compare cumulative coverage and off-list items.
- Read the revision aloud for naturalness.
- Ask a target learner to read or listen and explain the meaning.
- Restore any technical precision lost during simplification.
- Record the final lexical profile for future material design.
Part XLIX — A Curriculum Lexical Audit
A school or course can profile a representative sample of texts across months or years. The aim is not to produce a giant spreadsheet for its own sake. The aim is to see whether learners receive a coherent progression of lexical demand.
114. High-frequency foundation
Early materials should provide enough high-frequency vocabulary for learners to build fluency and meaning. If basic bands remain unstable, advanced text exposure becomes expensive.
115. Mid-frequency bridge
Learners need systematic exposure beyond K1–K2. A curriculum that never introduces mid-frequency vocabulary leaves students unprepared for authentic secondary, university and adult reading.
116. Academic layer
Academic vocabulary should recur across subjects rather than appear in one isolated English unit. Profiling can show where words such as analyse, factor, relevant, establish actually appear.
117. Technical layer
Each subject needs planned technical vocabulary. Profiling helps identify the size and recurrence of that layer so teachers can preteach and recycle rather than improvise.
118. Reading progression
The lexical step between one year and the next should be challenging but survivable. Profiling can detect sudden cliffs that deserve bridging materials.
Part L — The Mastery Club Profiling Loop
Use the eduKate learning loop at text level: Read → Diagnose → Prioritise → Repair → Practise → Connect → Perform → Review.
- Read: inspect the text and learner task before running statistics.
- Diagnose: profile vocabulary and identify where lexical demand sits.
- Prioritise: choose the bands, academic words or technical terms most relevant to actual comprehension.
- Repair: preteach, gloss, simplify or build domain knowledge where necessary.
- Practise: let learners encounter and retrieve the target vocabulary.
- Connect: reuse words across texts, subjects and phrase patterns.
- Perform: return to authentic reading or writing without excessive support.
- Review: compare observed performance with the profiler’s prediction and update future text selection.
Part LI — Profile Before You Simplify, Profile After You Simplify
This principle prevents two opposite mistakes. Without profiling first, teachers may simplify the wrong words and leave the real lexical bottleneck untouched. Without profiling after, they may assume the revision became more accessible when it merely became shorter. The before/after profile is a quality-control step, not a substitute for learner testing.
Part LII — The Mastery Club Standard for a Good Lexical Profile
- The text analysed is exactly the text the learner will encounter.
- The profiler and list version are recorded.
- The counting unit is stated.
- Proper nouns and compounds are handled consistently.
- Percentages are interpreted alongside actual word lists.
- Cumulative bands are reported where useful.
- Academic and technical layers are separated when relevant.
- Off-list items are manually inspected.
- Results are compared only with methodologically compatible profiles.
- Profile data is combined with learner evidence.
- No band is treated as inherently good or bad.
- The final decision serves comprehension, learning or writing quality rather than the metric itself.
Part LIII — Frequency-Band Interpretation Bank
Very high K1 token share
What it may mean: The text relies heavily on the most frequent vocabulary. Interpretive caution: Often supports accessibility, but can still contain hard syntax, abstract concepts, idioms and technical senses of common words.
Moderate K1 with strong K2
What it may mean: The text moves beyond the most basic core but remains relatively general. Interpretive caution: Common in mature general prose and intermediate learning materials.
Large K3–K5 layer
What it may mean: Mid-frequency vocabulary plays a substantial role. Interpretive caution: May signal advanced general reading, broad academic vocabulary or transitional difficulty for intermediate learners.
Large K6–K10 layer
What it may mean: A long lower-frequency tail appears. Interpretive caution: Could reflect specialist subject matter, literary vocabulary, advanced prose or unnecessary rarity.
Large K10+ layer
What it may mean: Very low-frequency vocabulary appears frequently. Interpretive caution: Usually requires expert, literary or highly specialised interpretation; inspect the actual items.
Large academic-list share
What it may mean: Formal cross-disciplinary vocabulary is prominent. Interpretive caution: Expected in academic prose; useful for school/university vocabulary planning.
Large technical-list share
What it may mean: Domain-specific terminology carries the text. Interpretive caution: General vocabulary scores may understate the importance of subject knowledge.
Large off-list share
What it may mean: Many tokens do not match the chosen framework. Interpretive caution: Could reflect names, errors, abbreviations, compounds, new words, technical terms or genuinely rare items.
Small off-list but high difficulty
What it may mean: Form-frequency profile looks accessible while comprehension remains hard. Interpretive caution: Investigate specialist senses, syntax, background knowledge, discourse and multiword language.
High K1 in advanced text
What it may mean: Common vocabulary is doing complex intellectual work. Interpretive caution: Frequency should not be confused with conceptual simplicity.
Low K1 in easy specialist text for experts
What it may mean: Specialist vocabulary is locally familiar to the audience. Interpretive caution: Audience expertise changes effective difficulty.
High K1 plus repeated technical terms
What it may mean: General frame is simple around a few specialist concepts. Interpretive caution: Good candidate for pre-teaching a compact Tier 3 set.
High type diversity but ordinary bands
What it may mean: Many different common words are used. Interpretive caution: Lexical diversity is high without necessarily increasing sophistication.
Low type diversity but advanced bands
What it may mean: A smaller set of advanced words repeats. Interpretive caution: Could reflect focused technical writing rather than limited vocabulary.
High academic share in learner writing
What it may mean: Learner uses many academic-list words. Interpretive caution: Check accuracy, collocation and whether the task naturally calls for them.
Falling K1 over time in comparable essays
What it may mean: More vocabulary beyond the highest-frequency band appears. Interpretive caution: Potential productive growth if accuracy and task comparability are maintained.
Stable K1 over time
What it may mean: Band distribution changes little. Interpretive caution: Vocabulary development may still occur in depth, phraseology, precision or comprehension.
High off-list in multilingual writing
What it may mean: Unmatched forms are frequent. Interpretive caution: Separate L1 items, names, spelling errors and legitimate rare vocabulary before interpretation.
High off-list in technology text
What it may mean: New product or computing terms are frequent. Interpretive caution: Reference lists may lag innovation; update or use a technical list.
High K1 in spoken transcript
What it may mean: Most tokens are common. Interpretive caution: Expected in speech; complexity may come from discourse structure and interaction rather than rare words.
Part LIV — Off-List Diagnosis Bank
| Off-list cause | Example | What to do |
|---|---|---|
| Proper noun | Maya, Singapore, OpenAI, Shakespeare | Decide whether entity knowledge matters; do not automatically count as difficult vocabulary. |
| Acronym | DNA, GDP, API, MRT | Expand and classify by domain; frequency lists may not recognise them. |
| Technical term | mitosis, subpoena, latency | Consider a custom technical list and teach with the concept. |
| Spelling error | enviroment, goverment | Correct if lexical range is the target; retain separately for form analysis. |
| Hyphenated compound | evidence-based, long-term | Check tokenisation; components may be high-frequency. |
| Closed compound | smartphone, cybersecurity | List age may lag language change. |
| Foreign-language item | bonjour, kampung, zeitgeist depending corpus | Determine whether borrowing or code-switching is intentional. |
| Neologism | deepfake, doomscroll | Modern usage may postdate the reference list. |
| Abbreviation | dept, approx | Normalise if the research question is vocabulary rather than orthographic convention. |
| Symbol string | CO2, H2O, x² | Usually remove or analyse separately for lexical profiling. |
| Markup noise | HTML class names, menu labels | Clean the source text before profiling. |
| Archaic word | thou, hath | Rarity may be stylistic/historical rather than general vocabulary weakness. |
| Dialect/regional word | local variety item | General corpora may underrepresent regional vocabulary. |
| Inflected/derived mismatch | rare form not linked by tool | Check lemma/family settings. |
| Named technical product | specific drug/device/model | May combine entity knowledge with domain vocabulary. |
Part LV — Text Comparison Mini-Lab
Two Primary science passages
Profile situation: Passage A has 91% K1–K2 and 4% technical vocabulary. Passage B has 85% K1–K2 and 8% technical vocabulary. Interpretation: B is probably lexically more demanding, but check whether the technical terms repeat and whether they were already taught.
Two Secondary history passages
Profile situation: Both reach 94% by K4. One has 6% proper nouns; the other has 6% abstract academic vocabulary. Interpretation: The raw cumulative profile hides different learning jobs: entity knowledge versus academic language.
Two learner essays
Profile situation: Both contain 86% K1. Essay A uses accurate K3–K5 words; Essay B uses many incorrect rare synonyms. Interpretation: Profiles are similar, but lexical control is not. Human evaluation separates sophistication from error.
Original and simplified article
Profile situation: Original reaches 95% cumulative coverage at K6; simplified reaches it at K3. Interpretation: The revision reduced lexical rarity, but confirm that technical concepts and natural phrasing survived.
News and textbook text
Profile situation: Both have 80% K1. News has names as off-list; textbook has technical terms. Interpretation: Same K1 does not mean same type of difficulty.
Lecture transcript and research paper
Profile situation: Lecture has 91% K1, paper 68% K1. Interpretation: The paper is lexically denser in lower-frequency vocabulary, but both may communicate equally advanced content.
General-health leaflet and medical guideline
Profile situation: Leaflet reaches 97% by K3; guideline needs K8 plus technical list. Interpretation: Audience design is visible in the profile.
Two AI rewrites
Profile situation: Rewrite A reduces K5+ but deletes key technical terms; Rewrite B keeps terms and simplifies surrounding words. Interpretation: Rewrite B is educationally superior despite slightly higher lexical demand.
Two graded readers
Profile situation: Both show 96% K1–K2; one has twenty unique off-list names, the other two repeated new vocabulary items. Interpretation: Type counts and recurrence make the learning burden different.
Two textbook chapters
Profile situation: Chapter A has a gradual mid-frequency profile; Chapter B contains a sudden technical cluster. Interpretation: Chapter B may need pre-teaching even if whole-chapter percentages are similar.
Part LVI — Profiling Questions by User
Teacher
- Which words in this text sit beyond the students’ likely vocabulary?
- Which are broadly useful academic words?
- Which are unavoidable technical terms?
- How many unfamiliar items are repeated enough to deserve deep teaching?
- Can I preserve the concept while simplifying nonessential rarity?
Student
- Why does this text feel hard?
- Which unknown words are worth learning first?
- Does my writing rely only on a small high-frequency core?
- Are my unusual words accurate and natural?
- Has my lexical range changed across comparable writing samples?
Parent
- Is this book lexically suitable for independent reading?
- Are the difficult words mostly technical terms that can be explained quickly?
- Would an easier version preserve enjoyment and comprehension?
- Which broadly useful words are worth discussing after reading?
Curriculum designer
- Does lexical demand progress gradually across years?
- Where do academic words recur across subjects?
- Which technical words are introduced but never recycled?
- Are there vocabulary cliffs between levels?
- Are simplified materials exposing learners to enough mid-frequency vocabulary for future growth?
Researcher
- What construct does the profile represent?
- What unit and list framework fit that construct?
- How will genre and text length be controlled?
- How will spelling, names and compounds be handled?
- What other measures are needed alongside the profile?
Part LVII — Canonical Map Inside the Vocabulary Ecosystem
| Question | Canonical owner |
|---|---|
| Broad vocabulary system | Vocabulary — The Mastery Club |
| How many words a learner knows | Vocabulary | Vocabulary Size |
| How many words vs how well known | Vocabulary Breadth and Vocabulary Depth |
| Vocabulary variety in a sample | Lexical Diversity |
| Text vocabulary distribution | This article — Lexical Profiling |
| Reader-known proportion of a text | Vocabulary | Lexical Coverage |
| How quickly known words can be retrieved | Vocabulary | Lexical Access |
| What knowing one word involves | Vocabulary | Word Knowledge |
| How vocabulary is organised in memory | Vocabulary | Mental Lexicon |
| Single words vs multiword units | Vocabulary | Lexical Chunks and Phrase Frames |
| How to select vocabulary for teaching | Tier 1, Tier 2 and Tier 3 Vocabulary |
The map prevents cannibalisation. Lexical Profiling owns the question “What vocabulary bands and list categories does this text contain?” Lexical Coverage owns “How much of this text does this reader know?” Vocabulary Size owns “How many words does the learner know?” Lexical Diversity owns “How varied is the vocabulary?” These are related but not interchangeable.
Part LVIII — Manual Profiling Exercise: Build a Profile Without Software
A manual exercise helps learners understand what software is doing. Take a short 100-word paragraph. Print or paste a frequency list beside it. Classify every running token into K1, K2, K3+ or off-list. Tally the counts. Then divide each band count by 100 to obtain percentages. The exercise is slow, but it exposes the assumptions hidden inside automated profiling.
119. Step 1 — Preserve the original text
Keep one untouched copy. Manual editing can accidentally change punctuation, words or forms. A preserved original lets you check disputed classifications later.
120. Step 2 — Mark every token
Do not classify only content words. Articles, auxiliaries and repeated function words contribute heavily to K1 token share. The profile describes the whole running text unless the method explicitly excludes categories.
121. Step 3 — Resolve ambiguous forms
If a word can belong to more than one lemma or part of speech, decide how the chosen list handles it. A surface-form exercise may ignore the distinction; a parsed system may not.
122. Step 4 — Separate off-list causes
Create subcategories: proper noun, spelling error, abbreviation, technical term, foreign item, unknown cause. This turns a blunt off-list count into an interpretable diagnosis.
123. Step 5 — Calculate token and type profiles
Token bands reveal reading-stream frequency. Type bands reveal lexical variety. Comparing both shows whether a band is represented by many different words or a few repeated ones.
124. Step 6 — Calculate cumulative coverage
Add K1+K2, then K1+K2+K3, and continue. Plotting the cumulative line can reveal how quickly the text becomes covered by increasingly broad vocabulary knowledge.
Part LIX — Worked 100-Word Example
Imagine a 100-token informational paragraph with the following classification: 82 K1 tokens, 8 K2 tokens, 4 K3 tokens, 3 academic-list tokens and 3 off-list tokens. Because the paragraph is exactly 100 tokens, the percentages are numerically identical to the counts: 82%, 8%, 4%, 3% and 3%.
The first interpretation is descriptive: high-frequency vocabulary dominates. The second step is cumulative: K1+K2 = 90%. If the academic-list items overlap with another category under the selected system, the accounting method must be clarified; list frameworks do not always create mutually exclusive bins in the same way. The third step is diagnostic: inspect the three off-list words. If they are all proper nouns, lexical rarity is lower than the raw off-list percentage suggests.
Now imagine that the four K3 tokens are all repetitions of one word family, while the three academic tokens are three distinct items. The text contains fewer unique K3 learning targets than the token percentage alone implies. Type and family counts therefore refine the instructional picture.
Part LX — Worked 1,000-Word Example
Consider a 1,000-token Secondary science chapter: K1 760 tokens, K2 90, K3–K5 60, academic-list 40, technical-list 35 and off-list 15. The general high-frequency core is 85%. Mid-frequency vocabulary adds 6%. Academic and technical layers together account for 7.5%, and off-list contributes 1.5%.
A weak interpretation would say the chapter is “15% difficult vocabulary.” A stronger interpretation inspects the forty academic and thirty-five technical tokens. If twenty-five technical tokens come from only five repeating terms, the technical learning set may be manageable. If the academic tokens include maintain, significant, factor, indicate, these are cross-curricular words worth deeper teaching. If most off-list tokens are species names, they may require entity recognition rather than vocabulary study.
Part LXI — Troubleshooting a Vocabulary Profiler
| Problem | Likely cause | Repair |
|---|---|---|
| Everything appears off-list | Wrong language/list selected, formatting corruption, or tokenisation failure. | Check input encoding, list framework and whether the text includes markup. |
| Proper nouns dominate output | Names are not being recategorised. | Use the profiler’s proper-noun options or inspect separately. |
| Common derived words appear off-list | The framework may be lemma-based or missing the derivation. | Check family/lemma settings and list version. |
| Hyphenated words behave strangely | Tokenizer is splitting or preserving hyphens inconsistently. | Normalise hyphens or use the same policy across texts. |
| Percentages do not total 100 | Categories may overlap, be displayed separately, or some tokens may be excluded. | Read the tool documentation and output definitions. |
| Two runs give different totals | Input changed, preprocessing changed or software settings differ. | Save the cleaned text and record settings. |
| Academic words seem too low | Selected academic list may not match the text genre or school level. | Try a more suitable academic/school framework. |
| Technical text looks impossibly rare | No domain list is being used. | Add a custom technical list or interpret off-list words manually. |
| Learner writing looks more advanced after spelling correction | Errors were previously counted as off-list. | Report both lexical profile and form accuracy separately. |
| Short text percentages swing dramatically | Sample is too small. | Use longer or multiple comparable samples before drawing conclusions. |
Part LXII — 25 Comparative Profiling Scenarios
A1 vs A2 graded readers
Comparison: A1: 96% K1–K2. A2: 91% K1–K2 with more K3–K5. Interpretation: A2 gives more lexical stretch. Choose based on learner vocabulary and reading purpose.
Primary science vs Primary story
Comparison: Science has more technical vocabulary; story has more low-frequency descriptive adjectives. Interpretation: Different difficulty types require different support.
News vs textbook
Comparison: News off-list is names; textbook off-list is subject terms. Interpretation: Off-list share alone cannot rank difficulty.
Two student narratives
Comparison: One uses mostly K1 with vivid collocations; the other uses more K5 words awkwardly. Interpretation: The first may demonstrate stronger lexical control despite a ‘simpler’ profile.
Two student arguments
Comparison: Both use similar bands; one has much higher lexical diversity. Interpretation: Frequency profile and diversity measure different properties.
Original vs simplified science text
Comparison: Simplified version raises K1 and preserves five technical terms. Interpretation: Good simplification if comprehension improves and concepts remain precise.
Original vs over-simplified science text
Comparison: K1 rises dramatically but technical terms are removed. Interpretation: Lexical ease improved at the cost of subject learning.
Lecture vs article
Comparison: Lecture is K1-heavy; article has more academic and K5+ vocabulary. Interpretation: Mode influences lexical distribution.
Professional email vs policy report
Comparison: Email is highly K1–K2; report is academic and technical. Interpretation: Audience and genre explain the difference.
General learner vs expert reader
Comparison: Same technical text profile for both. Interpretation: Profile is identical; lexical coverage and comprehension differ because learner knowledge differs.
British corpus vs American corpus profile
Comparison: Some word ranks shift. Interpretation: Corpus variety changes band assignment.
Old list vs updated list
Comparison: New technology terms move from off-list to recognised bands. Interpretation: List age changes interpretation.
Family-based vs lemma-based profile
Comparison: Family version shows lower lexical demand. Interpretation: Grouping derivatives reduces counted units.
Type-based vs token-based profile
Comparison: Type profile shows many K4 words; token profile is dominated by K1. Interpretation: A small number of frequent function words dominate running text.
Short paragraph vs full chapter
Comparison: Short paragraph has volatile percentages. Interpretation: Longer samples generally provide more stable estimates.
Literature excerpt vs complete novel
Comparison: Excerpt looks unusually rare due to one stylistic scene. Interpretation: Local profile may not represent the whole work.
Text with glossary vs without glossary
Comparison: Same lexical profile. Interpretation: Support changes reader experience without changing the text distribution.
Text before preteaching vs after preteaching
Comparison: Same lexical profile. Interpretation: Learner coverage changes even though the text profile does not.
AI text prompt A vs B
Comparison: Both ask for Grade 7; outputs differ strongly in K5+ vocabulary. Interpretation: Prompt labels do not guarantee lexical calibration.
Student draft before revision vs after revision
Comparison: Later draft replaces vague common verbs with accurate mid-frequency verbs. Interpretation: Potential improvement if appropriateness and meaning are stronger.
Student draft after thesaurus overuse
Comparison: K10+ increases sharply. Interpretation: Profile suggests rarity, but quality may decline.
Technical text with repeated term vs many unique terms
Comparison: Same technical token percentage. Interpretation: Unique learning burden differs sharply.
Text with proper-noun cluster vs technical cluster
Comparison: Same off-list percentage. Interpretation: Entity knowledge and conceptual vocabulary are different demands.
Expository vs narrative text
Comparison: Expository has more academic vocabulary; narrative more dialogue and K1 grammar. Interpretation: Genre should guide expectations.
Two curricula
Comparison: Curriculum A gradually increases K3–K6 exposure; B jumps suddenly in one year. Interpretation: Profiling reveals a potential lexical cliff in B.
Part LXIII — A 30-Day Profiling Practice Plan
- Day 1: profile a simple news paragraph and inspect off-list items.
- Day 2: profile a children’s story and compare proper nouns.
- Day 3: profile a science explainer.
- Day 4: compare token and type views.
- Day 5: compare family and lemma frameworks.
- Day 6: profile your own 300-word writing sample.
- Day 7: inspect every K4+ word manually.
- Day 8: profile a spoken transcript.
- Day 9: compare spoken and written versions of the same topic.
- Day 10: profile a technical text with a custom list.
- Day 11: calculate cumulative coverage manually.
- Day 12: clean a web article before profiling.
- Day 13: test how proper-noun handling changes percentages.
- Day 14: test how spelling correction changes a learner profile.
- Day 15: compare two graded readers.
- Day 16: compare an original and simplified passage.
- Day 17: identify Tier 2 candidates from a profile.
- Day 18: identify Tier 3 candidates from a domain text.
- Day 19: inspect multiword expressions missed by the profiler.
- Day 20: compare profile data with actual learner comprehension.
- Day 21: profile a research abstract.
- Day 22: profile a policy report.
- Day 23: profile an AI-generated educational text.
- Day 24: edit the AI text and re-profile.
- Day 25: profile two learner essays of the same genre.
- Day 26: evaluate whether lower-frequency words are accurate.
- Day 27: build a small curriculum vocabulary map.
- Day 28: identify a lexical cliff across two school levels.
- Day 29: write your profiling method in reproducible form.
- Day 30: explain what the profile can and cannot claim without looking at this guide.
Part LXIV — The Final Operating Principles
- Always record the profiler, list framework and version.
- Always state the counting unit.
- Never compare incompatible frameworks casually.
- Inspect off-list words manually.
- Treat proper nouns, technical terms and spelling errors separately when needed.
- Use token and type information for different questions.
- Use cumulative K-levels to understand the long tail of demand.
- Remember that frequency is corpus-dependent.
- Remember that common forms can carry technical senses.
- Remember that single-word profiling misses multiword meaning.
- Do not equate rarity with quality.
- Do not equate high-frequency vocabulary with intellectual simplicity.
- Combine profiling with coverage, comprehension and learner evidence.
- Use profiling to improve teaching decisions, not merely to generate numbers.
- Re-profile after substantial text adaptation.
- Prefer reproducible methods over undocumented intuition.
- Use the profile as a map, not a verdict.
Conclusion — A Lexical Profile Makes the Invisible Distribution Visible
Every text has a vocabulary shape. Some texts are built almost entirely from the highest-frequency bands. Others add a long mid-frequency tail, an academic layer, a technical layer or a cloud of names and off-list items. Lexical profiling makes that shape visible.
The power of the method lies in comparison and diagnosis. Teachers can calibrate reading materials, curriculum designers can detect lexical cliffs, students can inspect their productive range, and researchers can quantify aspects of lexical richness. But the method only works well when its assumptions remain visible: corpus, list, counting unit, text preparation, proper-noun handling and genre.
The most important distinction is the one between the profile of the text and the knowledge of the reader. A profiler can tell you where words sit in a frequency framework. It cannot know which words this learner knows, what concepts they understand, whether a phrase is idiomatic, or whether a rare word is precisely the right word.
Return to Vocabulary — The Mastery Club for the broad vocabulary map, Vocabulary | Lexical Coverage for reader-known text proportion, Vocabulary | Vocabulary Size for learner breadth, and Lexical Diversity for vocabulary variety. This page remains the canonical eduKate owner for lexical profiling, vocabulary profiling, Lexical Frequency Profiles, K-levels and the measurement of vocabulary demands in text.
Appendix A — Choosing a Profiler for the Job
| Goal | Profiler choice | Why |
|---|---|---|
| I want a simple historic LFP-style profile | Use a Classic K1/K2/academic/off-list framework. | Good for teaching the original logic and for comparison with classic research. |
| I want fine-grained general-English bands | Use BNC/COCA 1–25k or another multi-band system. | Good for advanced reading and text comparison. |
| I care about productive writing | Prefer lemma/form-sensitive analysis where appropriate. | Avoid assuming derivational-family knowledge automatically. |
| I care about receptive reading | Family-based bands can be useful if morphological assumptions fit the learners. | Remember that family knowledge is not guaranteed. |
| I am analysing university prose | Use general bands plus an academic list. | Academic list membership helps separate formal vocabulary from general low-frequency vocabulary. |
| I am analysing lectures | Use a spoken-academic framework where available. | Written academic lists may misrepresent spoken patterns. |
| I am analysing school textbooks | Use age-appropriate school lists plus general bands. | University lists can be too remote from school vocabulary. |
| I am analysing technical material | Add a custom technical list or domain corpus. | Otherwise technical words may collapse into off-list. |
| I am comparing texts over time | Use the same profiler, list version, unit and preprocessing rules. | Methodological consistency matters more than chasing every new tool. |
| I am checking AI-generated educational text | Use a modern general-frequency framework and manually inspect technical terms. | Verify the output; do not trust the requested grade label alone. |
Appendix B — K-Band Planning Table
| Band | Approximate band position | Instructional interpretation |
|---|---|---|
| K1 | 1–1000 band under a 1k system | High-frequency core; strong priority for broad comprehension. |
| K2 | 1001–2000 band under a 1k system | High-frequency core; strong priority for broad comprehension. |
| K3 | 2001–3000 band under a 1k system | Mid-frequency bridge; increasingly important for independent advanced reading. |
| K4 | 3001–4000 band under a 1k system | Mid-frequency bridge; increasingly important for independent advanced reading. |
| K5 | 4001–5000 band under a 1k system | Mid-frequency bridge; increasingly important for independent advanced reading. |
| K6 | 5001–6000 band under a 1k system | Lower-frequency general vocabulary; value depends strongly on reading goals and genre. |
| K7 | 6001–7000 band under a 1k system | Lower-frequency general vocabulary; value depends strongly on reading goals and genre. |
| K8 | 7001–8000 band under a 1k system | Lower-frequency general vocabulary; value depends strongly on reading goals and genre. |
| K9 | 8001–9000 band under a 1k system | Lower-frequency general vocabulary; value depends strongly on reading goals and genre. |
| K10 | 9001–10000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
| K11 | 10001–11000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
| K12 | 11001–12000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
| K13 | 12001–13000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
| K14 | 13001–14000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
| K15 | 14001–15000 band under a 1k system | Low-frequency tail; often specialist, literary or advanced, requiring selective rather than blanket study. |
The ranges above describe the idea of successive 1,000-unit bands, not a universal master list. Actual members depend on the chosen framework, corpus and counting unit. K5 in one system is not automatically the same lexical set as K5 in another.
Appendix C — Worked Cumulative Coverage Table
| Text | K1 | K2 | K3 | K4 | K5 | Other/off | Reading |
|---|---|---|---|---|---|---|---|
| Text 1 | 88 | 7 | 2 | 1 | 1 | 1 | 95% by K2; strongly high-frequency. |
| Text 2 | 80 | 8 | 5 | 3 | 2 | 2 | 93% by K3; moderate mid-frequency demand. |
| Text 3 | 72 | 8 | 6 | 4 | 4 | 6 | 80% by K2; long lexical tail. |
| Text 4 | 91 | 4 | 1 | 1 | 1 | 2 | 95% by K2; inspect off-list names. |
| Text 5 | 76 | 8 | 5 | 3 | 2 | 6 | 84% by K2; lower-frequency and off-list layer needs inspection. |
| Text 6 | 69 | 7 | 6 | 5 | 4 | 9 | 76% by K2; highly specialised or advanced. |
| Text 7 | 84 | 8 | 3 | 2 | 1 | 2 | 92% by K2; small K3–K5 layer. |
| Text 8 | 78 | 6 | 5 | 4 | 3 | 4 | 84% by K2; more extended mid-frequency vocabulary. |
| Text 9 | 87 | 5 | 3 | 1 | 1 | 3 | 92% by K2; off-list may matter more than higher bands. |
| Text 10 | 74 | 7 | 5 | 4 | 4 | 6 | 81% by K2; requires broad vocabulary or strong domain knowledge. |
These are illustrative profiles, not empirical benchmark texts. Their purpose is to train cumulative interpretation: the useful question is how quickly the profile reaches high cumulative percentages, what remains outside the common bands, and whether the remaining vocabulary is general, academic, technical, named or erroneous.
Appendix D — Reproducibility Checklist
- Save the exact raw text.
- Save the cleaned analysis text.
- Record the date of analysis.
- Record the profiler name and version.
- Record the list framework.
- Record whether units are families, lemmas or forms.
- Record tokenisation rules if they were changed.
- Record how contractions were handled.
- Record how hyphens were handled.
- Record how proper nouns were handled.
- Record how numbers and symbols were handled.
- Record spelling-normalisation decisions.
- Record any custom vocabulary list.
- Record total tokens after cleaning.
- Record total types if relevant.
- Record band percentages.
- Record cumulative band percentages.
- Save the off-list items.
- Save technical-list matches.
- Record any manual recategorisation.
- Use identical settings for comparison texts.
- Report genre and task.
- Report text length.
- Report learner population if making pedagogical claims.
- State clearly what the profile cannot measure.
Appendix E — 40 Profiling Questions to Ask Before You Trust the Output
- 1. What exact text was analysed?
- 2. Was the text cleaned?
- 3. Did navigation or metadata remain?
- 4. Which profiler produced the output?
- 5. Which list framework was selected?
- 6. Which corpus underlies the list?
- 7. Are the units families or lemmas?
- 8. How are inflections handled?
- 9. How are derivations handled?
- 10. How are proper nouns handled?
- 11. How are abbreviations handled?
- 12. How are numbers handled?
- 13. How are hyphenated compounds handled?
- 14. Were spelling errors corrected?
- 15. Is the sample long enough?
- 16. Is the genre comparable to other samples?
- 17. Does the framework match the learner age?
- 18. Does it match spoken or written mode?
- 19. Are technical terms separated?
- 20. Are academic words separated?
- 21. What proportion is K1?
- 22. What proportion is K2?
- 23. What cumulative band reaches 95%?
- 24. What cumulative band reaches 98%?
- 25. What are the most frequent higher-band words?
- 26. How many higher-band types are unique?
- 27. How many repeat?
- 28. What actually appears off-list?
- 29. Are names inflating the off-list share?
- 30. Are errors inflating it?
- 31. Are common forms being used in specialist senses?
- 32. Are multiword expressions invisible to the profiler?
- 33. Does learner vocabulary data support the interpretation?
- 34. Does observed comprehension support it?
- 35. Does the profile fit the purpose of the text?
- 36. Would a different corpus change the conclusion?
- 37. Would a lemma profile change the conclusion?
- 38. Would a family profile change the conclusion?
- 39. Are you describing the text or judging the learner?
- 40. What decision will this analysis actually change?
Appendix F — The Profiler-to-Teaching Translation
If K1–K2 is very high
Focus less on broad general vocabulary and more on syntax, concepts, discourse and the small number of high-impact unknown items. A high-frequency profile suggests the lexical core is not the only likely bottleneck.
If K3–K5 is substantial
Identify recurring mid-frequency words with cross-text value. These can be strong candidates for Tier 2 teaching, especially if they appear across subjects or units.
If academic-list share is high
Teach the words in disciplinary contexts and phrase patterns. Academic vocabulary should support reasoning rather than become a detached prestige list.
If technical-list share is high
Build domain knowledge. Use diagrams, examples, procedures and concept maps. Technical vocabulary should remain attached to the discipline that gives it meaning.
If off-list share is high
Diagnose before teaching. Separate names, errors, compounds, neologisms and technical terms. Off-list is an investigation queue, not a vocabulary syllabus.
If a few words repeat many times
Prioritise them. Learning one recurring family can raise effective coverage of the entire unit quickly.
If many different low-frequency words occur once
Consider whether the text is appropriate for independent reading. The unique learning burden may be too high even if the total token percentage looks manageable.
Appendix G — The Mastery Club Final Standard
A complete lexical profile should let a reader answer five questions: What vocabulary layers does this text contain? Which layers are likely to matter for this learner? Which words are repeated enough to deserve teaching? Which apparent difficulties are only profiling artefacts? What action should change because of the analysis? If the profile cannot answer those questions, the analysis is not yet educationally complete.
The method should end in a decision: keep the text, scaffold it, simplify selected wording, preteach vocabulary, build background knowledge, choose a different text, or use the text for a different reading mode. Numbers without action are inventory. Profiling becomes teaching only when it changes what happens next.
Appendix H — Reliability, Validity and Decision Error
A lexical profile looks objective because it produces percentages, but the quality of the conclusion depends on the quality of the measurement design. Reliability asks whether similar analysis conditions produce stable results. Validity asks whether the profile actually represents the construct we care about. Decision error asks what happens if we use the profile to choose a text, judge a writer or design an intervention and the interpretation is wrong.
Reliability problem 1 — Text length
Short samples produce volatile percentages. In a fifty-word sample, one unusual word changes the profile by two percentage points. In a 1,000-word sample, the same single token changes the percentage by only 0.1. This does not mean longer is always better, but it means sample size must be considered before comparing small differences.
Reliability problem 2 — Genre
The same writer uses different vocabulary in a personal narrative, laboratory report and argumentative essay. If genre changes, the lexical profile may change even when the learner’s underlying vocabulary ability has not. Longitudinal comparisons should therefore control task and genre where possible.
Reliability problem 3 — Topic
Topic changes vocabulary. A familiar topic may activate specialist words the learner already knows; an unfamiliar topic may constrain productive range. A single profile can reflect topic knowledge as much as general lexical capability.
Reliability problem 4 — Editing support
A timed classroom essay and a polished take-home draft are not equivalent samples. Dictionaries, spellcheckers, AI tools and teacher feedback can alter lexical choice. Record production conditions when comparing profiles.
Validity problem 1 — Frequency is not knowledge
A word’s corpus frequency does not prove that a learner knows or does not know it. Personal interests, profession, multilingual background and prior study all shape lexical knowledge. Frequency bands estimate population-level exposure patterns, not individual mastery.
Validity problem 2 — Frequency is not difficulty
Common words can be hard in specialised senses; rare words can be easy for experts. Difficulty also depends on morphology, context, syntax and conceptual knowledge. A lexical profile captures only one dimension of text demand.
Validity problem 3 — Frequency is not quality
A well-written sentence can use entirely common vocabulary. An awkward sentence can use rare vocabulary. If the research question is writing quality, lexical profiling should be paired with measures of accuracy, appropriateness, cohesion and argument.
Validity problem 4 — Family membership is not full knowledge
A family-based profiler assumes a meaningful relationship among derived forms. Learners may not know every member or sense. Family profiling is a modelling decision, not direct evidence of morphological mastery.
Decision error 1 — Rejecting a good text
A teacher may reject a science text because it contains many low-frequency technical terms, even though those terms are exactly what the unit must teach. The better decision may be to keep the text and scaffold the technical layer.
Decision error 2 — Accepting a difficult text
A text can appear easy because most forms are K1–K2 while using common words in technical senses or complex relationships. A high-frequency profile should never be the sole criterion for independent reading.
Decision error 3 — Rewarding thesaurus writing
If a teacher rewards higher-band percentages without judging naturalness, students may replace precise common words with obscure alternatives. This trains metric gaming rather than vocabulary development.
Decision error 4 — Misdiagnosing learner weakness
If comprehension fails, the profile may tempt us to blame vocabulary. The real bottleneck could be syntax, background knowledge, inference or attention. Use profile data as one diagnostic channel, not the whole diagnosis.
Appendix I — A Validity Triangle
A trustworthy educational interpretation should align three kinds of evidence. Text evidence comes from the lexical profile: bands, academic words, technical terms and off-list items. Learner evidence comes from vocabulary knowledge, reading behaviour, writing samples and observed errors. Task evidence comes from what the learner must actually do: skim, study, infer, write, solve, explain or perform under time pressure.
When all three point toward the same problem, the diagnosis is stronger. If the profile shows a large technical layer, the learner does not know the technical terms, and the task requires detailed subject comprehension, vocabulary is a credible bottleneck. If the profile looks easy but the learner still fails, look elsewhere.
Appendix J — Reporting a Lexical Profile Properly
A professional report should be reproducible. A minimal methods sentence might state that a cleaned text was profiled using a specified tool and list framework, with words counted as families or lemmas, proper nouns handled in a stated way, and percentages reported by K-level. If results are used to compare texts, the same preprocessing and framework should be used across all texts.
A stronger report also states the genre, token count, custom list use, off-list inspection procedure and limitations. If the analysis informs teaching, include learner evidence rather than presenting text statistics alone.
Appendix K — Twenty Final Checks Before Publication or Classroom Use
- 1. Have I recorded the profiler version?
- 2. Have I recorded the list framework?
- 3. Do I know whether units are families, lemmas or forms?
- 4. Is the input text clean?
- 5. Are proper nouns handled consistently?
- 6. Are numbers and symbols treated appropriately?
- 7. Have spelling errors been separated from lexical rarity?
- 8. Have I inspected the actual off-list words?
- 9. Have I inspected repeated K3+ words?
- 10. Have I separated academic from technical vocabulary where relevant?
- 11. Am I comparing texts of compatible genre and length?
- 12. Am I using the same framework for every comparison?
- 13. Have I checked for specialist senses of common words?
- 14. Have I considered multiword expressions the profiler misses?
- 15. Have I checked actual learner vocabulary?
- 16. Have I checked observed comprehension or writing quality?
- 17. Have I avoided treating rare words as automatically good?
- 18. Have I avoided treating K1 words as automatically easy?
- 19. Do I know what decision the profile will change?
- 20. Can another person reproduce the analysis from my notes?
Final Synthesis — Lexical Profiling as a Measurement Discipline
Lexical profiling is most powerful when treated as measurement rather than decoration. The profiler does not merely colour-code words. It operationalises assumptions about corpus frequency, lexical units, list membership and text composition. Good analysis makes those assumptions visible; weak analysis hides them behind percentages.
For educators, the payoff is practical. A profile can identify where a text’s lexical demand comes from, distinguish general vocabulary from technical terminology, reveal whether a revision actually reduced rarity, and show which vocabulary layers recur across a curriculum. For learners, it can make reading difficulty and productive vocabulary more visible. For researchers, it provides a reproducible quantitative lens on lexical distribution.
The final rule is simple: profile the text, inspect the words, check the learner, respect the task. Only then turn the numbers into an educational decision.
Appendix L — Profile-to-Action Matrix
| Profile pattern | Likely interpretation | Best next move |
|---|---|---|
| K1–K2 above 95%, low off-list | Lexical core is probably accessible for many learners. | Focus on syntax, concepts, discourse and the few high-impact unfamiliar items rather than adding a large vocabulary list. |
| K1–K2 below 85%, long K3–K8 tail | Mid-frequency vocabulary is carrying a substantial share of the text. | Consider pre-teaching recurring general words, selecting an easier text for independent reading, or using guided reading. |
| High academic-list share | The text relies heavily on general academic vocabulary. | Teach recurring cross-curricular words through morphology, collocation and reasoning functions. |
| High technical-list share | Specialist terminology is central. | Teach the domain concepts explicitly; preserve the technical words rather than replacing them with vague language. |
| High off-list caused by proper nouns | Entity density is inflating the profile. | Separate names from lexical difficulty; build background knowledge if the entities matter. |
| High off-list caused by spelling errors | The profile is contaminated by form errors. | Correct for lexical-range analysis and report spelling separately. |
| High off-list caused by new technology words | Reference list may lag current usage. | Use a newer corpus/list or create a technical/custom category. |
| Low-frequency words repeat heavily | A small number of families create much of the demand. | Prioritise those repeated items; one teaching event can improve access to many later tokens. |
| Many unique low-frequency words occur once | The lexical learning burden is distributed widely. | The text may be unsuitable for independent reading even if the total low-frequency token percentage seems moderate. |
| Common words carry technical senses | Frequency profile understates conceptual difficulty. | Teach the specialised meanings and subject models explicitly. |
| Student writing has very high K1 | Productive range may be narrow, or the task may naturally invite common words. | Check genre, precision and repetition before recommending rarer vocabulary. |
| Student writing has more K5+ vocabulary | Range may be expanding. | Verify correctness, collocation and register before treating it as improvement. |
| Simplified text raises K1 but removes technical terms | Lexical profile improved while conceptual fidelity declined. | Restore essential terminology and simplify surrounding wording instead. |
| Simplified text raises K1 while retaining key technical words | Accessibility probably improved without losing the concept. | Test with target learners and preserve natural phraseology. |
| Profile predicts difficulty but learners perform well | Learners may have strong domain knowledge, cognate support or prior exposure. | Update the learner model; general frequency was overpredicting demand. |
| Profile predicts ease but learners struggle | The bottleneck may be syntax, inference, background knowledge, phrases or technical senses. | Do not prescribe more vocabulary automatically; diagnose the first weak link. |
| Two texts differ only slightly in K1 percentage | Difference may be trivial. | Inspect actual higher-band words, text length and learner performance before choosing between them. |
| Two texts have similar cumulative coverage but different off-list types | The same percentage hides different lexical jobs. | Choose support based on whether off-list items are entities, technical terms, errors or rare general words. |
| Curriculum shows a sudden K3–K6 jump | Learners face a lexical cliff. | Insert bridge texts, targeted vocabulary and repeated mid-frequency exposure before the jump. |
| Curriculum remains almost entirely K1–K2 for years | Learners may be underexposed to authentic lexical growth. | Introduce controlled mid-frequency and academic vocabulary so later authentic reading does not arrive as a shock. |
The Final Decision Rule
A lexical profile should never end with “this text has 82% K1 vocabulary.” The analysis becomes educational only when the number is translated into a decision. Keep the text, scaffold the text, simplify selected wording, preteach recurring words, build the subject concepts, or change the reading mode. If the profile does not change what you do, it is descriptive information without instructional consequence.
The strongest use of profiling therefore combines three layers: the statistical distribution of the text, the actual word list behind the distribution, and the real learner response. Those three together turn frequency bands into a usable model of vocabulary demand.
Appendix M — Minimal Methods Note for Reusable Profiling
When lexical profiling is used repeatedly, the method should be written once in a reusable form. A minimal methods note can prevent silent drift between analyses. Record the input source, date, genre, token count, cleaning rules, profiler version, list framework, lexical unit, proper-noun policy, spelling-normalisation policy, technical-list use and the exact outputs you intend to compare.
Why method drift matters
Suppose Text A is profiled with BNC/COCA families and Text B with an NGSL lemma framework. The percentages may look comparable because both outputs contain frequency bands, but the underlying categories are different. The comparison is not merely imperfect; it answers a different question. Method drift can also occur when one text has names recategorised and another leaves them off-list, or when learner spelling is corrected in one sample but not another.
Why version drift matters
Profiling tools evolve. Lists are updated, tokenisation can change, proper-noun handling can improve and new list frames can be added. Current Lextutor interfaces explicitly show ongoing version changes. If an analysis will be reproduced months later, recording the tool version protects comparability.
Why purpose must come first
A profile for choosing an independent-reading text is not identical to a profile for evaluating productive lexical richness. Receptive reading may justify family-based analysis; productive writing may benefit from lemma-sensitive analysis. The method should therefore begin with the construct: what exactly are you trying to know?
Why human inspection remains mandatory
No output can interpret every unmatched form correctly. A profiler does not know whether Jordan is a person, country or brand; whether charge is legal, electrical or financial; whether deepfake is unfamiliar to this learner; or whether maintain was used naturally. Manual inspection is not a failure of automation. It is part of valid lexical analysis.
A Reusable Reporting Sentence
A clear report can state that the text was profiled using a named tool and version, under a named frequency-list framework, with words counted as families or lemmas, proper nouns handled by a stated rule, obvious spelling errors treated according to a stated policy, and results reported as token percentages by K-level with manual inspection of off-list items. That one sentence makes the analysis far more interpretable and reproducible.
Final Principle
Do not compare numbers whose measurement systems you cannot explain. Lexical profiling is valuable because it makes vocabulary distribution measurable. Its credibility depends on remembering that every measurement is produced by a method. Keep the method visible, and the profile becomes a reliable tool for reasoning rather than a decorative statistic.
