Vocabulary keyness is a corpus-linguistic way to find keywords and distinctive vocabulary by comparing a target corpus with a relevant reference corpus. Instead of asking only which words are frequent, keyness asks which words occur unusually often or unusually rarely relative to a comparison norm. If you are searching for keyness, keyword analysis, corpus keywords, distinctive vocabulary, target corpus, reference corpus, word frequency, log-likelihood, text dispersion, corpus linguistics or how to find important words in a text collection, this guide builds the complete system.
Vocabulary frequency and vocabulary keyness are not the same measurement. A very frequent function word may occur at similar rates everywhere and therefore tell us little about what makes a target corpus distinctive. A less frequent content word can become highly key if it appears disproportionately often in the target material. Keyness is therefore useful for identifying the aboutness, style, register, preoccupations and lexical profile of texts, subjects, communities, genres and learner corpora—provided the comparison corpus is chosen carefully.
This Vocabulary | Keyness article is an additive specialist owner under The Mastery Club / Master Vocabulary router. It extends Vocabulary | Corpus Linguistics without replacing it, and connects to Lexical Profiling, word frequency, dispersion, collocation, concordance and vocabulary teaching. The clean ownership question here is: which lexical items distinguish this target collection from an appropriate reference?
The 50-Second Router
- Frequency: how often a word appears in one corpus.
- Keyness: how unusually frequent or infrequent a word is in a target corpus compared with a reference corpus.
- Target corpus: the texts you want to characterise.
- Reference corpus: the comparison set that defines what counts as ordinary or expected.
- Positive keyword: an item unusually frequent in the target relative to the reference.
- Negative keyword: an item unusually infrequent in the target relative to the reference.
- Important caution: keyword lists are comparison-dependent; change the reference corpus and the result can change.
- Best practice: combine keyness statistics with dispersion, concordance reading, context and a clearly stated research question.
1. What Keyness Means
Keyness is a comparative property. A word becomes key not simply because it appears many times, but because its rate in a target corpus differs meaningfully from its rate in a reference corpus.
Corpus tools operationalise this comparison with statistical measures. Different methods rank items differently, so keyness should be treated as an analytical lens rather than an intrinsic property permanently attached to a word.
The same lexical item can be highly key in one comparison, ordinary in another and negatively key in a third.
For vocabulary education, this teaches a powerful idea: what characterises a text or domain emerges through contrast.
2. Frequency Versus Keyness
A frequency list orders words inside one corpus. A keyword list asks which words stand out relative to another corpus.
The most frequent words in English are often grammatical function words. They matter enormously for language, but they may not identify a topic because they occur everywhere.
A domain term can be less frequent in raw count yet more distinctive. In a climate corpus, terms such as emissions, carbon or renewable may separate the target from general language.
Therefore, teachers should not equate ‘most frequent’ with ‘most characteristic’ or ‘most important for this domain’.
3. The Target Corpus
The target corpus is the text collection whose vocabulary you want to understand. It might contain essays, news stories, scientific papers, speeches, forum posts, learner writing or one author’s work.
Good target-corpus design begins with a research question. If the question is about Grade 9 science vocabulary, the corpus should genuinely represent that material rather than an arbitrary set of webpages.
Metadata matters: date, genre, author, audience, region and topic can all influence lexical patterns.
Keyness cannot rescue a badly defined target. Clear sampling is part of lexical analysis.
4. The Reference Corpus
The reference corpus supplies the comparison norm. It may be a large balanced general corpus or a closely matched corpus that differs on one dimension of interest.
Reference choice is not a technical afterthought. It determines what counts as unusual. Comparing medical articles with general English highlights medical vocabulary; comparing two medical specialties highlights subtler specialist differences.
A reference that differs simultaneously in topic, era, genre and audience can generate a keyword list full of confounds.
The strongest analysis explains why the reference is appropriate for the exact question.
5. Positive and Negative Keywords
Positive keywords occur more than expected in the target relative to the reference. Negative keywords occur less than expected.
Positive keyness often receives more attention because it highlights distinctive vocabulary, but negative keyness can reveal absences, avoided forms or different stylistic choices.
Neither label means good or bad. Positive and negative describe direction of relative frequency.
Students should always state the direction and comparison instead of saying vaguely that a word ‘has high keyness’.
6. Keyness and Aboutness
Keyword lists can reveal what a text collection is strongly about because content words associated with recurring subjects rise above the reference norm.
A corpus of vaccination debates, for example, may show topic terms that are not merely frequent but unusually concentrated relative to a matched reference.
Aboutness is not automatic interpretation. A keyword tells us where to look; concordance lines show how the item is actually used.
The analytical sequence should therefore move from ranking to context rather than treating the keyword list as the final answer.
7. Keyness and Style
Function words, pronouns, modal verbs and other grammatical items can also become key. These may indicate stylistic, interpersonal or register differences rather than topic.
A target corpus may overuse first-person pronouns, certainty markers or particular conjunctions compared with the reference.
Such patterns can be educationally valuable because vocabulary choice includes stance and discourse organisation, not only content nouns.
Interpret stylistic keywords carefully and connect them to actual examples in context.
8. Normalised Frequency
Target and reference corpora are often different sizes, so raw counts cannot usually be compared directly. Normalised frequency expresses occurrences relative to a common number of tokens.
For example, frequencies per million words allow rates to be compared even when one corpus is much larger.
Keyness statistics typically work with counts while accounting for corpus sizes mathematically, but normalized rates remain useful for human interpretation.
Students should distinguish raw frequency, normalised frequency and keyness score instead of mixing them.
9. Statistical Measures of Keyness
Corpus tools offer several keyness measures, including log-likelihood, chi-square, Fisher-type tests, log ratio, Bayesian measures and newer dispersion-aware approaches.
A significance-oriented measure asks whether a difference is unlikely under a null model; an effect-size measure asks how large the relative difference is. These are related but not identical questions.
Large corpora can make tiny differences statistically significant, so effect size and practical importance deserve attention.
A responsible article or lesson reports the method rather than presenting a ranked list as method-free truth.
10. Log-Likelihood
Log-likelihood is widely used in corpus linguistics to compare frequencies across corpora. It produces larger values as the evidence for a difference increases under its model.
The number itself does not express a familiar everyday unit. It is a ranking/statistical device that must be interpreted alongside counts, corpus sizes and ideally effect size.
Students do not need to calculate the formula by hand to understand the logic: observed distributions are compared with expected distributions.
The key pedagogical lesson is that distinctiveness is a comparison of distributions, not just a count.
11. Effect Size and Log Ratio
Effect-size measures such as log ratio help describe the magnitude and direction of relative frequency difference.
A word can have strong statistical evidence because corpora are huge while the actual proportional difference remains small. Conversely, a rare term can show a large ratio but rest on very few occurrences.
Combining evidence helps avoid overclaiming.
For vocabulary teaching, prioritize terms that are both meaningfully distinctive and sufficiently recurrent or important to the domain.
12. Dispersion
Dispersion asks how broadly an item is distributed across texts or sections. A word occurring one hundred times in one document behaves differently from a word occurring four times in each of twenty-five documents.
Traditional frequency-based keyness can be influenced by bursty local repetition. Dispersion-aware approaches try to account for distribution.
The educational interpretation is straightforward: a broadly dispersed keyword is more likely to represent the domain or register as a whole than a word produced by one anomalous document.
Always inspect range or dispersion before turning a keyword list into a teaching list.
13. Text Dispersion Keyness
Text dispersion keyness gives weight to how widely a feature occurs across texts, not only how many total tokens it contributes.
Recent corpus work has used dispersion-aware methods to distinguish corpora while reducing domination by a small number of highly repetitive texts.
No method eliminates the need for judgement; dispersion thresholds and corpus structure still matter.
The method is especially useful when the unit of analysis is genuinely a collection of texts rather than one long undifferentiated stream.
14. Keywords Need Concordance Analysis
A keyword list identifies candidate features. Concordance lines reveal meanings, collocations, grammatical functions, evaluative patterns and repeated contexts.
One form can have several senses, so raw keyness may combine different uses. Concordance reading separates them.
This is also where surprising keywords become interpretable. A common word can be key because one domain repeatedly uses a specialised sense.
The workflow should therefore be quantitative discovery followed by qualitative interpretation.
15. Keyness and Collocation
After identifying a keyword, collocation analysis can show which words cluster around it. This reveals local phraseology and semantic framing.
Two corpora may both contain the same keyword but surround it with different verbs, adjectives or stance markers.
For domain vocabulary teaching, keyword-plus-collocation is more useful than the isolated keyword because learners need conventional use.
The learner can build a lexical profile: keyword, definition, key collocates, grammar, register and representative examples.
16. Keyness and Lexical Profiling
Lexical profiling asks what frequency bands or vocabulary levels make up a text. Keyness asks what distinguishes one corpus from another. The two methods answer complementary questions.
A text can contain mostly high-frequency vocabulary yet still have a small set of highly distinctive keywords.
Conversely, a text can be lexically difficult because of many low-frequency words even if those words are not distinctive relative to another specialist corpus.
Keeping these measures separate prevents analytical confusion and improves curriculum decisions.
17. Keyness and Domain Vocabulary
Keyness is useful for discovering candidate domain vocabulary because specialist terms often occur disproportionately in subject-specific texts.
But not every key item should be taught. Proper names, formatting artifacts, temporary event terms or corpus-specific noise can rise in rankings.
Human filtering remains necessary: relevance, generalisability, learner level, curricular importance and transfer potential matter.
The best teaching list is therefore corpus-informed, not corpus-obedient.
18. Keyness and Learner Corpora
Researchers can compare learner writing with another learner corpus or a reference corpus to identify distinctive lexical and grammatical patterns.
The interpretation depends heavily on corpus matching. Differences may reflect task, proficiency, first language, genre or topic rather than a single learner characteristic.
Keyword analysis can generate useful hypotheses, but causal claims require additional evidence.
For classroom use, the safest application is descriptive: identify recurring patterns, inspect examples and design targeted practice.
19. Keyness and Comparing Genres
A school science explanation, newspaper report and personal narrative may use different distinctive vocabulary even when discussing related content.
Comparing genres can reveal verbs of evidence, stance markers, reporting expressions, technical nouns and discourse organisers that define each communicative job.
This helps students learn vocabulary as genre-bound action rather than as isolated synonym lists.
A genre vocabulary lesson should combine keywords with examples of the rhetorical functions they serve.
20. Keyness and Time
A corpus from one period can be compared with a matched corpus from another period to identify emerging and declining vocabulary.
Temporal comparisons can reveal technological change, social concerns, new terminology and shifts in style.
However, spelling conventions, corpus sources and genre balance must be controlled as far as possible.
A key word across time is evidence of changed relative usage, not automatically evidence of changed meaning.
21. Keyness and Communities
Communities of practice can develop distinctive vocabulary and phraseology. Comparing their language with a relevant reference may surface terms, abbreviations and recurrent expressions.
Distinctiveness can reflect shared expertise, identity, routines or institutional roles.
Ethical interpretation matters when community data concerns real people. Keyword patterns should not be used to stereotype individuals.
For education, focus on domain literacy: what vocabulary newcomers need to understand participation in the community.
22. Stopwords and Function Words
Some analyses remove very common function words before keyword extraction; others keep them because grammatical words can reveal style and discourse.
There is no universally correct stopword policy. The choice depends on whether the question concerns topic vocabulary, style, grammar or all three.
Automatically removing function words can discard meaningful patterns; automatically keeping everything can bury content terms.
State the policy and inspect the effect instead of treating software defaults as methodology.
23. Lemmas, Word Families and Tokenisation
Keyword results depend on what counts as the same item. Analysing surface forms separately may split analyse, analyses, analysed, analysing; lemmatisation can combine related inflections.
Word-family grouping goes further and can obscure distinctions if used carelessly. Tokenisation rules also affect punctuation, contractions, hyphens and multiword expressions.
The unit should match the research and teaching question.
For domain teaching, a lemma-level keyword may be useful, followed by explicit teaching of the forms learners actually encounter.
24. Multiword Keyness
Distinctiveness can occur at phrase level, not just single words. N-grams, clusters, lexical bundles and key multiword expressions can characterise a genre or community.
This connects keyness with phraseology. A conventional phrase may be more revealing than any component word alone.
Phrase-level analysis also reduces the danger of interpreting a word without its recurrent frame.
Advanced vocabulary work should therefore inspect both key words and key phrases.
25. The End State: Comparative Vocabulary Intelligence
Keyness turns a pile of text into a comparative vocabulary question. It does not decide importance automatically; it shows where target language departs from a chosen norm.
The strongest workflow defines the target, justifies the reference, computes keyness, checks dispersion, reads concordances, analyses collocation, filters noise and interprets with domain knowledge.
For learners, the result can become a precise map of vocabulary worth noticing in a genre or subject.
For researchers and teachers, the deeper lesson is methodological humility: distinctive always means distinctive relative to something.
Keyness Laboratory: 120 Comparative Corpus Cases
The cases below are illustrative teaching comparisons rather than claimed empirical results. They train the reasoning that must occur before and after software produces a keyword list. Students predict, compare, inspect and interpret instead of treating a ranking as automatic truth.
Case 1: climate-policy reports vs general news — target/reference fit
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 2: climate-policy reports vs general news — frequency versus keyness
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 3: climate-policy reports vs general news — dispersion
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 4: climate-policy reports vs general news — concordance
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 5: climate-policy reports vs general news — collocation
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 6: climate-policy reports vs general news — teaching value
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 7: climate-policy reports vs general news — method
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 8: climate-policy reports vs general news — interpretation
Hypothetical comparison: target = climate-policy reports; reference = general news. Candidate vocabulary might include emissions, carbon, renewable, mitigation, targets, transition. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 9: medical research abstracts vs general academic prose — target/reference fit
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 10: medical research abstracts vs general academic prose — frequency versus keyness
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 11: medical research abstracts vs general academic prose — dispersion
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 12: medical research abstracts vs general academic prose — concordance
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 13: medical research abstracts vs general academic prose — collocation
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 14: medical research abstracts vs general academic prose — teaching value
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 15: medical research abstracts vs general academic prose — method
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 16: medical research abstracts vs general academic prose — interpretation
Hypothetical comparison: target = medical research abstracts; reference = general academic prose. Candidate vocabulary might include patients, clinical, treatment, outcomes, diagnosis, trial. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 17: school science textbooks vs general school prose — target/reference fit
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 18: school science textbooks vs general school prose — frequency versus keyness
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 19: school science textbooks vs general school prose — dispersion
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 20: school science textbooks vs general school prose — concordance
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 21: school science textbooks vs general school prose — collocation
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 22: school science textbooks vs general school prose — teaching value
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 23: school science textbooks vs general school prose — method
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 24: school science textbooks vs general school prose — interpretation
Hypothetical comparison: target = school science textbooks; reference = general school prose. Candidate vocabulary might include experiment, evidence, variable, energy, system, process. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 25: mathematics explanations vs general school prose — target/reference fit
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 26: mathematics explanations vs general school prose — frequency versus keyness
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 27: mathematics explanations vs general school prose — dispersion
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 28: mathematics explanations vs general school prose — concordance
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 29: mathematics explanations vs general school prose — collocation
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 30: mathematics explanations vs general school prose — teaching value
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 31: mathematics explanations vs general school prose — method
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 32: mathematics explanations vs general school prose — interpretation
Hypothetical comparison: target = mathematics explanations; reference = general school prose. Candidate vocabulary might include equation, ratio, function, calculate, value, represent. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 33: legal judgments vs general formal prose — target/reference fit
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 34: legal judgments vs general formal prose — frequency versus keyness
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 35: legal judgments vs general formal prose — dispersion
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 36: legal judgments vs general formal prose — concordance
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 37: legal judgments vs general formal prose — collocation
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 38: legal judgments vs general formal prose — teaching value
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 39: legal judgments vs general formal prose — method
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 40: legal judgments vs general formal prose — interpretation
Hypothetical comparison: target = legal judgments; reference = general formal prose. Candidate vocabulary might include court, claim, evidence, statute, appeal, liability. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 41: financial reports vs general business news — target/reference fit
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 42: financial reports vs general business news — frequency versus keyness
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 43: financial reports vs general business news — dispersion
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 44: financial reports vs general business news — concordance
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 45: financial reports vs general business news — collocation
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 46: financial reports vs general business news — teaching value
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 47: financial reports vs general business news — method
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 48: financial reports vs general business news — interpretation
Hypothetical comparison: target = financial reports; reference = general business news. Candidate vocabulary might include revenue, assets, liability, quarter, earnings, cash. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 49: technology support pages vs general web prose — target/reference fit
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 50: technology support pages vs general web prose — frequency versus keyness
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 51: technology support pages vs general web prose — dispersion
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 52: technology support pages vs general web prose — concordance
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 53: technology support pages vs general web prose — collocation
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 54: technology support pages vs general web prose — teaching value
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 55: technology support pages vs general web prose — method
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 56: technology support pages vs general web prose — interpretation
Hypothetical comparison: target = technology support pages; reference = general web prose. Candidate vocabulary might include device, settings, update, account, password, install. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 57: travel guides vs general informational prose — target/reference fit
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 58: travel guides vs general informational prose — frequency versus keyness
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 59: travel guides vs general informational prose — dispersion
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 60: travel guides vs general informational prose — concordance
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 61: travel guides vs general informational prose — collocation
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 62: travel guides vs general informational prose — teaching value
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 63: travel guides vs general informational prose — method
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 64: travel guides vs general informational prose — interpretation
Hypothetical comparison: target = travel guides; reference = general informational prose. Candidate vocabulary might include hotel, route, museum, station, booking, district. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 65: restaurant reviews vs general lifestyle prose — target/reference fit
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 66: restaurant reviews vs general lifestyle prose — frequency versus keyness
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 67: restaurant reviews vs general lifestyle prose — dispersion
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 68: restaurant reviews vs general lifestyle prose — concordance
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 69: restaurant reviews vs general lifestyle prose — collocation
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 70: restaurant reviews vs general lifestyle prose — teaching value
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 71: restaurant reviews vs general lifestyle prose — method
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 72: restaurant reviews vs general lifestyle prose — interpretation
Hypothetical comparison: target = restaurant reviews; reference = general lifestyle prose. Candidate vocabulary might include dish, menu, flavour, service, chef, portion. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 73: sports reports vs general news — target/reference fit
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 74: sports reports vs general news — frequency versus keyness
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 75: sports reports vs general news — dispersion
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 76: sports reports vs general news — concordance
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 77: sports reports vs general news — collocation
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 78: sports reports vs general news — teaching value
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 79: sports reports vs general news — method
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 80: sports reports vs general news — interpretation
Hypothetical comparison: target = sports reports; reference = general news. Candidate vocabulary might include match, score, coach, season, player, league. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 81: music criticism vs general arts reviews — target/reference fit
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 82: music criticism vs general arts reviews — frequency versus keyness
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 83: music criticism vs general arts reviews — dispersion
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 84: music criticism vs general arts reviews — concordance
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 85: music criticism vs general arts reviews — collocation
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 86: music criticism vs general arts reviews — teaching value
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 87: music criticism vs general arts reviews — method
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 88: music criticism vs general arts reviews — interpretation
Hypothetical comparison: target = music criticism; reference = general arts reviews. Candidate vocabulary might include performance, tempo, melody, orchestra, rhythm, composition. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 89: literary analysis essays vs general student essays — target/reference fit
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 90: literary analysis essays vs general student essays — frequency versus keyness
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 91: literary analysis essays vs general student essays — dispersion
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 92: literary analysis essays vs general student essays — concordance
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 93: literary analysis essays vs general student essays — collocation
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 94: literary analysis essays vs general student essays — teaching value
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 95: literary analysis essays vs general student essays — method
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 96: literary analysis essays vs general student essays — interpretation
Hypothetical comparison: target = literary analysis essays; reference = general student essays. Candidate vocabulary might include character, theme, narrator, symbol, motif, conflict. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 97: argumentative essays vs personal narratives — target/reference fit
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 98: argumentative essays vs personal narratives — frequency versus keyness
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 99: argumentative essays vs personal narratives — dispersion
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 100: argumentative essays vs personal narratives — concordance
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 101: argumentative essays vs personal narratives — collocation
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 102: argumentative essays vs personal narratives — teaching value
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 103: argumentative essays vs personal narratives — method
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 104: argumentative essays vs personal narratives — interpretation
Hypothetical comparison: target = argumentative essays; reference = personal narratives. Candidate vocabulary might include evidence, therefore, claim, however, issue, because. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 105: news reports vs personal blogs — target/reference fit
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 106: news reports vs personal blogs — frequency versus keyness
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 107: news reports vs personal blogs — dispersion
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 108: news reports vs personal blogs — concordance
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 109: news reports vs personal blogs — collocation
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 110: news reports vs personal blogs — teaching value
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 111: news reports vs personal blogs — method
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 112: news reports vs personal blogs — interpretation
Hypothetical comparison: target = news reports; reference = personal blogs. Candidate vocabulary might include reported, according, officials, statement, incident, investigation. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 113: research articles vs popular science articles — target/reference fit
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 114: research articles vs popular science articles — frequency versus keyness
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 115: research articles vs popular science articles — dispersion
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 116: research articles vs popular science articles — concordance
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 117: research articles vs popular science articles — collocation
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 118: research articles vs popular science articles — teaching value
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 119: research articles vs popular science articles — method
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
Case 120: research articles vs popular science articles — interpretation
Hypothetical comparison: target = research articles; reference = popular science articles. Candidate vocabulary might include method, results, participants, significant, analysis, data. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. Do not invent a statistical result. Instead state what evidence would be needed before calling any candidate key.
Next, design the verification step. Record target frequency, reference frequency, corpus sizes, normalized rates, range across texts and the selected keyness/effect-size measures. Read concordance lines for the strongest candidates. Then decide whether each item indicates topic, register, style, corpus artifact or sampling imbalance. This makes keyword analysis a transparent method rather than a black-box list.
A 10-Step Keyness Workflow
1. Define the question
Write one sentence stating exactly what lexical distinctiveness you want to investigate. Document the decision so another analyst could understand how the vocabulary list was produced.
2. Build the target corpus
Collect texts that validly represent the population, genre, period or domain of interest. Document the decision so another analyst could understand how the vocabulary list was produced.
3. Choose the reference
Select a corpus that creates the comparison required by the question, not merely the largest corpus available. Document the decision so another analyst could understand how the vocabulary list was produced.
4. Clean and tokenise
Decide how to handle metadata, punctuation, casing, lemmas, contractions and non-language artifacts. Document the decision so another analyst could understand how the vocabulary list was produced.
5. Generate frequency lists
Inspect raw and normalised frequencies in both corpora before statistical ranking. Document the decision so another analyst could understand how the vocabulary list was produced.
6. Compute keyness
Use an explicit method such as log-likelihood and consider an effect-size measure. Document the decision so another analyst could understand how the vocabulary list was produced.
7. Check dispersion
Verify whether apparent keywords are broadly distributed or produced by a few texts. Document the decision so another analyst could understand how the vocabulary list was produced.
8. Read concordances
Inspect senses, contexts, collocations, stance and phraseology. Document the decision so another analyst could understand how the vocabulary list was produced.
9. Filter and interpret
Remove artifacts and classify meaningful items by topic, style, register or domain function. Document the decision so another analyst could understand how the vocabulary list was produced.
10. Validate teaching use
If building a vocabulary list, check learner level, transfer value, curriculum relevance and examples before publication. Document the decision so another analyst could understand how the vocabulary list was produced.
Keyness Misconceptions
“Keywords are just the most frequent words.”
No. Keywords are unusually frequent or infrequent relative to a reference; raw frequency and keyness answer different questions.
“A key word is objectively important everywhere.”
No. Keyness depends on the target/reference comparison and method.
“The biggest reference corpus is always best.”
No. Suitability to the research question matters more than size alone.
“A significant keyword is automatically educationally important.”
No. Statistical distinctiveness does not replace curricular judgement.
“One burst of repetition proves a domain keyword.”
Not necessarily. Check text dispersion and range.
“Function words should always be removed.”
No. They may carry stylistic and grammatical distinctiveness, depending on the question.
“A keyword list explains why the pattern exists.”
No. It identifies a difference; explanation needs concordance evidence, context and often other data.
“Changing the reference should not change the keywords.”
False. The reference defines the norm, so different comparisons can produce different results.
Research Anchors
- Cambridge Handbook of English Corpus Linguistics (2026) — Keyword Analysis
- Cambridge — Applying Corpus Linguistics to Illness and Healthcare, introduction to keywords and keyness
- Lancaster University — WordSmith Keyword Analysis
- Cambridge — Introducing the Chinese Learner English Corpus: text-dispersion keyness example
FAQ
What is keyness in corpus linguistics?
Keyness is a measure of how unusually frequent or infrequent an item is in a target corpus compared with a reference corpus.
What is a keyword?
In corpus analysis, a keyword is an item identified as statistically distinctive relative to a chosen reference, not simply a search-engine keyword or important-looking term.
What is a reference corpus?
It is the comparison corpus that supplies the expected or normal frequency pattern.
Why not just use a frequency list?
Frequency shows commonness inside one corpus; keyness shows comparative distinctiveness.
What is a positive keyword?
A word that is relatively overrepresented in the target corpus.
What is a negative keyword?
A word that is relatively underrepresented in the target corpus.
Does keyness prove a text is about a topic?
It can provide strong clues to aboutness, but interpretation should be checked with context and concordance evidence.
Why does dispersion matter?
A word repeated heavily in one text may look important in total counts without representing the corpus broadly.
Which keyness statistic is best?
There is no single method that is best for every question. Report the method and consider significance, effect size, dispersion and corpus design together.
Can keyness build vocabulary lists?
Yes, as one evidence source. Human filtering is still needed for relevance, level, transfer value and noise.
How Keyness Connects to the eduKate Vocabulary Ecosystem
Start with The Mastery Club. Use Corpus Linguistics for the broader data system, Lexical Profiling for frequency-band demands, and Word Frequency for general frequency. Keyness owns the comparative distinctiveness layer: target against reference.
For phrase-level patterns, continue to Vocabulary | Phraseology and Collocations and Word Partnerships. This allows a keyword to expand into its characteristic phrases rather than remaining an isolated token.
Final Synthesis
Keyness is comparative vocabulary intelligence. It asks not simply what appears, but what appears unusually often or unusually rarely against a justified norm. That small change in question turns frequency data into a method for discovering lexical distinctiveness.
The method is strongest when statistical ranking is followed by dispersion checks, concordance reading, collocation analysis and human interpretation. For education, the result is a disciplined way to identify candidate domain and genre vocabulary without pretending that software can decide curricular importance by itself.
Return to The Mastery Club vocabulary router for the complete vocabulary architecture.
Applied Workshop 1: Designing a primary science keyness study
Design a small study in which the target resembles climate-policy reports and the reference resembles general news. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 2: Designing a secondary science keyness study
Design a small study in which the target resembles medical research abstracts and the reference resembles general academic prose. Write three plausible concordance contexts and identify whether one form carries multiple senses or functions. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 3: Designing a mathematics keyness study
Design a small study in which the target resembles school science textbooks and the reference resembles general school prose. State what a report must disclose: corpus sizes, sampling, reference choice, keyness measure, thresholds and dispersion treatment. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 4: Designing a history keyness study
Design a small study in which the target resembles mathematics explanations and the reference resembles general school prose. Predict which candidate may be frequent in both corpora and which may be distinctive despite a smaller raw count. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 5: Designing a geography keyness study
Design a small study in which the target resembles legal judgments and the reference resembles general formal prose. Add likely lexical partners and explain how phrase-level evidence clarifies the keyword’s role. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 6: Designing a literature keyness study
Design a small study in which the target resembles financial reports and the reference resembles general business news. Separate descriptive evidence from causal claims and write one cautious conclusion that the data could support. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 7: Designing a news keyness study
Design a small study in which the target resembles technology support pages and the reference resembles general web prose. Ask whether the candidate appears across many texts or is concentrated in one document, and explain how that changes interpretation. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 8: Designing a academic writing keyness study
Design a small study in which the target resembles travel guides and the reference resembles general informational prose. Decide whether the keyword belongs in a learner list after filtering names, noise, temporary event terms and overly specialised items. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
Applied Workshop 9: Designing a oral transcripts keyness study
Design a small study in which the target resembles restaurant reviews and the reference resembles general lifestyle prose. Explain why the reference corpus is or is not a fair norm for the target and name one confound that could distort the comparison. List the sampling decisions that would make the comparison fair: period, genre, audience, text length, source and topic balance. Then predict three kinds of output separately—topic terms, stylistic items and possible artifacts—without claiming that any prediction is a measured result.
Create a reporting table specification with fields for raw target frequency, raw reference frequency, normalised rates, keyness statistic, effect size and text range. Explain what each column contributes. Add a concordance step so a keyword with several senses is not misinterpreted. If the target contains repeated templates or boilerplate, decide whether these should be removed or analysed as part of the register.
Finally convert the hypothetical keyword list into a teaching decision. Keep only items that are relevant to the learning goal, sufficiently representative, understandable at the learner’s level and transferable beyond one source. Add definitions, collocations and examples, then compare the resulting teaching list with a simple frequency list. Explain why the two lists should not be expected to match.
