DISTRIBUTIONAL SEMANTICS · DISTRIBUTIONAL HYPOTHESIS · CONTEXT · CO-OCCURRENCE · SEMANTIC VECTORS · COSINE SIMILARITY · CORPUS LINGUISTICS
Distributional semantics studies word meaning through patterns of linguistic distribution. Instead of beginning with a handcrafted definition, it asks where a word occurs, which words surround it, which grammatical relations it enters and how its contextual profile compares with other words.
The central distributional idea is that words used in similar contexts tend to share semantic properties. That principle underlies corpus-based semantic models, modern word embeddings and many AI systems. It also provides a powerful vocabulary-learning insight: repeated contextual company teaches learners what kind of word they are dealing with before every detail is explicitly defined.
This guide explains co-occurrence, context windows, vectors, cosine similarity, semantic neighbourhoods, corpus dependence, polysemy, contextual diversity, analogy, bias and grounding limits. Existing eduKateSG Semantic Relations, Mental Lexicon and Corpus Linguistics owners remain untouched.
Meaning leaves a statistical footprint in context.
2. Distributional semantics
Distributional semantics studies meaning through patterns of linguistic distribution. The central idea is that words used in similar contexts tend to have related meanings or functions.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
3. Distributional hypothesis
The distributional hypothesis is often summarised by the idea that a word can be partly known by the company it keeps. Context patterns provide evidence about semantics even without a handcrafted dictionary definition.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
4. Co-occurrence
Co-occurrence counts which words appear near one another. Repeated co-occurrence creates a statistical profile that can distinguish words occupying different semantic environments.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
5. Context windows
A distributional model needs a definition of context: nearby words, sentences, documents, syntactic dependencies or other structural units. Different windows capture different kinds of similarity.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
6. Bag-of-words models
Simple models may count nearby words without preserving order. Despite their simplicity, such representations can recover substantial semantic structure.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
7. Dependency-based contexts
Syntactic contexts can represent who does what to whom, often producing more functionally specific similarity than broad neighbouring-word windows.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
8. Vectors
Distributional representations encode a word as a vector whose dimensions summarise contextual behaviour. Similar vectors represent words with similar distributions.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
9. Cosine similarity
Cosine similarity is widely used to compare semantic vectors by their direction in a high-dimensional space rather than raw magnitude.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
10. Similarity versus relatedness
Distributional similarity can capture both category similarity and broader topical relatedness, depending on the model and context definition.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
11. Polysemy
A single vector may blur several senses of a polysemous word. Sense-aware models attempt to separate different contextual uses.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
12. Frequency effects
Very frequent words provide many observations but may also have broad, diffuse distributions. Rare words have sparse evidence and noisier representations.
For vocabulary learning, the important question is what contextual evidence this dimension preserves and what information it loses. Distributional evidence is powerful because it is systematic, but it is never the whole of meaning.
13. Contextual diversity
Words seen across varied contexts often develop more robust distributional profiles than words repeated in one narrow setting.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
14. Semantic neighbourhoods
Distributional spaces create neighbourhoods in which semantically or functionally similar words cluster together.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
15. Analogy and structure
Some embedding spaces show regular geometric patterns that support analogical relations, though these should not be treated as perfect logical rules.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
16. Corpus dependence
Distributional meaning depends on the corpus. A legal corpus, school corpus and social-media corpus can produce different semantic spaces.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
17. Bias
Because distributional models learn from language use, they can encode social stereotypes, exclusions and historical biases present in the data.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
18. Grounding limits
Distribution alone does not provide complete conceptual grounding. Words referring to perception, action and the physical world also depend on experience beyond text.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
19. Human learning
Human learners also exploit distributional information. Repeated contextual patterns help infer categories, relations and likely uses of new words.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
20. Vocabulary teaching
Distributional thinking encourages teachers to present words across meaningful contextual families rather than as isolated translations.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
21. Corpus learning
Concordance lines let learners inspect distribution directly: which verbs, nouns, adjectives and semantic classes repeatedly surround the target?
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
22. AI-era vocabulary
Modern language models depend heavily on contextual distribution. Understanding distributional semantics therefore helps learners distinguish statistical language knowledge from full world understanding.
A reliable interpretation should compare multiple contexts rather than one memorable sentence. Corpus composition, sense mixture and register can change the apparent semantic neighbourhood substantially.
23. Distributional casebook — Cases 1–20
1. doctor
Illustrative context profile: nurse,hospital,patient,clinic,medicine. Domain: healthcare contexts. Lesson: profession + institution.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
2. teacher
Illustrative context profile: student,classroom,lesson,school,teach. Domain: education contexts. Lesson: profession + activity.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
3. knife
Illustrative context profile: cut,sharp,blade,kitchen,slice. Domain: tool/action contexts. Lesson: instrument semantics.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
4. scissors
Illustrative context profile: cut,snip,paper,hair,sharp. Domain: tool/action contexts. Lesson: instrument semantics.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
5. rain
Illustrative context profile: cloud,wet,storm,weather,umbrella. Domain: weather contexts. Lesson: event semantics.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
6. snow
Illustrative context profile: cold,winter,ice,white,weather. Domain: weather contexts. Lesson: event/property semantics.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
7. bank
Illustrative context profile: money,loan,account,finance,interest. Domain: financial sense. Lesson: polysemy cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
8. bank
Illustrative context profile: river,shore,water,stream,flood. Domain: river-edge sense. Lesson: polysemy cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
9. cell
Illustrative context profile: biology,membrane,nucleus,tissue,organism. Domain: biological sense. Lesson: technical sense cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
10. cell
Illustrative context profile: prison,inmate,jail,locked,guard. Domain: prison sense. Lesson: polysemy cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
11. model
Illustrative context profile: data,predict,training,system,parameter. Domain: computational sense. Lesson: technical cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
12. model
Illustrative context profile: fashion,photograph,runway,agency,pose. Domain: person/fashion sense. Lesson: polysemy cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
13. run
Illustrative context profile: race,fast,walk,exercise,track. Domain: movement sense. Lesson: verb sense.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
14. run
Illustrative context profile: business,manage,company,operate,organisation. Domain: management sense. Lesson: verb sense.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
15. light
Illustrative context profile: bright,dark,lamp,shine,room. Domain: illumination sense. Lesson: noun/adjective cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
16. light
Illustrative context profile: weight,heavy,carry,thin,portable. Domain: weight sense. Lesson: polysemy cluster.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
17. strong
Illustrative context profile: powerful,weak,strength,muscle,force. Domain: physical evaluation. Lesson: paradigmatic similarity.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
18. strong
Illustrative context profile: evidence,argument,case,support,result. Domain: abstract academic use. Lesson: collocational extension.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
19. heavy
Illustrative context profile: weight,load,carry,light. Domain: physical sense. Lesson: adjective distribution.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
20. heavy
Illustrative context profile: rain,traffic,pressure,loss. Domain: abstract/collocational use. Lesson: semantic preference.
Neighbourhood test: list several words that appear in similar contexts and ask whether they are synonyms, category neighbours, topical associates or grammatical alternatives.
Sense test: inspect whether one semantic vector would merge distinct senses. If so, split concordance lines by sense before drawing conclusions.
Corpus test: predict how the profile would change in a specialist corpus, children’s corpus or social-media corpus.
Learning use: teach the target with several representative contexts so learners can infer the recurring distributional pattern rather than memorising one sentence.
24. Distributional casebook — Cases 21–40
21. evidence
Illustrative profile: support,claim,data,result,proof,study. Domain: academic reasoning. Lesson: argument vocabulary.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
22. hypothesis
Illustrative profile: test,theory,predict,experiment,evidence. Domain: science/research. Lesson: academic concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
23. algorithm
Illustrative profile: data,compute,model,search,process. Domain: computing. Lesson: technical concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
24. enzyme
Illustrative profile: protein,reaction,catalyst,cell,substrate. Domain: biology. Lesson: technical concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
25. inflation
Illustrative profile: prices,economy,rate,cost,central bank. Domain: economics. Lesson: domain concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
26. jurisdiction
Illustrative profile: court,law,authority,case,legal. Domain: law. Lesson: domain concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
27. tempo
Illustrative profile: music,beat,rhythm,speed,performance. Domain: music. Lesson: domain concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
28. offside
Illustrative profile: football,referee,goal,line,player. Domain: sport. Lesson: domain concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
29. selfie
Illustrative profile: photo,phone,camera,post,social media. Domain: digital culture. Lesson: modern lexicalisation.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
30. podcast
Illustrative profile: episode,audio,listen,host,show. Domain: media. Lesson: modern lexicalisation.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
31. happy
Illustrative profile: joy,smile,glad,pleased,sad. Domain: emotion. Lesson: semantic neighbourhood.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
32. angry
Illustrative profile: mad,furious,annoyed,rage,upset. Domain: emotion. Lesson: semantic neighbourhood.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
33. walk
Illustrative profile: run,stroll,road,feet,move. Domain: motion. Lesson: verb neighbourhood.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
34. whisper
Illustrative profile: speak,quiet,voice,soft,talk. Domain: speech manner. Lesson: troponymic relation.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
35. purchase
Illustrative profile: buy,price,customer,sale,payment. Domain: commerce. Lesson: formal/general overlap.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
36. buy
Illustrative profile: sell,money,shop,pay,price. Domain: commerce. Lesson: general register.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
37. child
Illustrative profile: parent,school,young,family,play. Domain: human/social. Lesson: life-stage concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
38. infant
Illustrative profile: baby,child,birth,mother,newborn. Domain: human/social. Lesson: more specific concept.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
39. vehicle
Illustrative profile: car,truck,bus,transport,road. Domain: category. Lesson: hypernymic neighbourhood.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
40. car
Illustrative profile: vehicle,drive,road,wheel,engine. Domain: category member. Lesson: hyponymic neighbourhood.
Window test: compare a narrow window of nearby words with a whole-document context. Narrow windows often capture local functional similarity; broad windows often capture topic.
Similarity test: choose two candidate neighbours and explain why cosine similarity might rank one closer even if human speakers judge the other more conceptually related.
Frequency test: consider whether rare contexts are underrepresented. Distributional models need enough observations to estimate patterns reliably.
Transfer: find a new word from the same domain and predict its distribution before checking corpus evidence.
25. Distributional casebook — Cases 41–60
41. fruit
Illustrative profile: apple,banana,orange,food,tree. Domain: category. Lesson: hypernymic neighbourhood.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
42. apple
Illustrative profile: fruit,red,tree,eat,juice. Domain: category member. Lesson: hyponymic neighbourhood.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
43. justice
Illustrative profile: law,fair,rights,court,equality. Domain: abstract social concept. Lesson: semantic field.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
44. freedom
Illustrative profile: rights,liberty,choice,independence,democracy. Domain: abstract social concept. Lesson: semantic field.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
45. data
Illustrative profile: analysis,information,collect,result,statistics. Domain: research/computing. Lesson: cross-domain term.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
46. information
Illustrative profile: data,knowledge,source,provide,detail. Domain: general abstract noun. Lesson: distribution contrast.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
47. energy
Illustrative profile: power,electricity,heat,renewable,consume. Domain: science/public. Lesson: cross-domain term.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
48. power
Illustrative profile: energy,control,authority,electricity,strength. Domain: polysemous general term. Lesson: multiple clusters.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
49. spring
Illustrative profile: season,flower,warm,summer,winter. Domain: season sense. Lesson: polysemy.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
50. spring
Illustrative profile: water,source,river,natural,flow. Domain: water-source sense. Lesson: polysemy.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
51. virus
Illustrative profile: infection,disease,cell,spread,immune. Domain: biology. Lesson: medical meaning.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
52. virus
Illustrative profile: computer,malware,file,security,infect. Domain: computing metaphor. Lesson: metaphorical sense.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
53. cloud
Illustrative profile: rain,sky,weather,storm,grey. Domain: weather sense. Lesson: literal sense.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
54. cloud
Illustrative profile: server,data,storage,computing,service. Domain: computing sense. Lesson: technical metaphor.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
55. root
Illustrative profile: plant,soil,tree,grow,stem. Domain: botanical sense. Lesson: literal sense.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
56. root
Illustrative profile: word,prefix,suffix,morphology,meaning. Domain: linguistic sense. Lesson: technical metaphor.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
57. field
Illustrative profile: farm,land,crop,grass,farmer. Domain: physical sense. Lesson: literal sense.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
58. field
Illustrative profile: research,study,discipline,academic,area. Domain: abstract domain sense. Lesson: metaphorical extension.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
59. frame
Illustrative profile: picture,wood,wall,photo,border. Domain: object sense. Lesson: literal sense.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
60. frame
Illustrative profile: scene,role,semantics,background,concept. Domain: linguistic/cognitive sense. Lesson: technical extension.
Polysemy audit: identify separate clusters in the context profile. Distinct clusters often correspond to senses, registers or domains.
Grounding audit: state what cannot be recovered from co-occurrence alone—sensory form, physical experience, causal mechanism or cultural practice.
Bias audit: ask whether corpus stereotypes could distort the neighbourhood. Distributional similarity reflects usage, including biased usage.
Teaching audit: preserve useful contextual patterns while explicitly correcting misleading associations.
26. A distributional-semantics learning protocol
1. doctor
Collect: gather authentic examples containing doctor. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as nurse,hospital,patient,clinic,medicine by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
2. teacher
Collect: gather authentic examples containing teacher. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as student,classroom,lesson,school,teach by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
3. knife
Collect: gather authentic examples containing knife. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as cut,sharp,blade,kitchen,slice by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
4. scissors
Collect: gather authentic examples containing scissors. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as cut,snip,paper,hair,sharp by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
5. rain
Collect: gather authentic examples containing rain. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as cloud,wet,storm,weather,umbrella by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
6. snow
Collect: gather authentic examples containing snow. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as cold,winter,ice,white,weather by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
7. bank
Collect: gather authentic examples containing bank. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as money,loan,account,finance,interest by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
8. bank
Collect: gather authentic examples containing bank. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as river,shore,water,stream,flood by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
9. cell
Collect: gather authentic examples containing cell. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as biology,membrane,nucleus,tissue,organism by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
10. cell
Collect: gather authentic examples containing cell. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as prison,inmate,jail,locked,guard by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
11. model
Collect: gather authentic examples containing model. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as data,predict,training,system,parameter by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
12. model
Collect: gather authentic examples containing model. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as fashion,photograph,runway,agency,pose by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
13. run
Collect: gather authentic examples containing run. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as race,fast,walk,exercise,track by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
14. run
Collect: gather authentic examples containing run. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as business,manage,company,operate,organisation by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
15. light
Collect: gather authentic examples containing light. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as bright,dark,lamp,shine,room by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
16. light
Collect: gather authentic examples containing light. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as weight,heavy,carry,thin,portable by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
17. strong
Collect: gather authentic examples containing strong. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as powerful,weak,strength,muscle,force by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
18. strong
Collect: gather authentic examples containing strong. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as evidence,argument,case,support,result by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
19. heavy
Collect: gather authentic examples containing heavy. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as weight,load,carry,light by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
20. heavy
Collect: gather authentic examples containing heavy. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as rain,traffic,pressure,loss by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
21. evidence
Collect: gather authentic examples containing evidence. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as support,claim,data,result,proof,study by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
22. hypothesis
Collect: gather authentic examples containing hypothesis. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as test,theory,predict,experiment,evidence by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
23. algorithm
Collect: gather authentic examples containing algorithm. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as data,compute,model,search,process by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
24. enzyme
Collect: gather authentic examples containing enzyme. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as protein,reaction,catalyst,cell,substrate by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
25. inflation
Collect: gather authentic examples containing inflation. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as prices,economy,rate,cost,central bank by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
26. jurisdiction
Collect: gather authentic examples containing jurisdiction. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as court,law,authority,case,legal by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
27. tempo
Collect: gather authentic examples containing tempo. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as music,beat,rhythm,speed,performance by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
28. offside
Collect: gather authentic examples containing offside. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as football,referee,goal,line,player by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
29. selfie
Collect: gather authentic examples containing selfie. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as photo,phone,camera,post,social media by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
30. podcast
Collect: gather authentic examples containing podcast. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as episode,audio,listen,host,show by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
31. happy
Collect: gather authentic examples containing happy. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as joy,smile,glad,pleased,sad by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
32. angry
Collect: gather authentic examples containing angry. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as mad,furious,annoyed,rage,upset by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
33. walk
Collect: gather authentic examples containing walk. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as run,stroll,road,feet,move by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
34. whisper
Collect: gather authentic examples containing whisper. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as speak,quiet,voice,soft,talk by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
35. purchase
Collect: gather authentic examples containing purchase. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as buy,price,customer,sale,payment by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
36. buy
Collect: gather authentic examples containing buy. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as sell,money,shop,pay,price by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
37. child
Collect: gather authentic examples containing child. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as parent,school,young,family,play by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
38. infant
Collect: gather authentic examples containing infant. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as baby,child,birth,mother,newborn by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
39. vehicle
Collect: gather authentic examples containing vehicle. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as car,truck,bus,transport,road by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
40. car
Collect: gather authentic examples containing car. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as vehicle,drive,road,wheel,engine by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
41. fruit
Collect: gather authentic examples containing fruit. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as apple,banana,orange,food,tree by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
42. apple
Collect: gather authentic examples containing apple. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as fruit,red,tree,eat,juice by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
43. justice
Collect: gather authentic examples containing justice. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as law,fair,rights,court,equality by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
44. freedom
Collect: gather authentic examples containing freedom. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as rights,liberty,choice,independence,democracy by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
45. data
Collect: gather authentic examples containing data. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as analysis,information,collect,result,statistics by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
46. information
Collect: gather authentic examples containing information. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as data,knowledge,source,provide,detail by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
47. energy
Collect: gather authentic examples containing energy. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as power,electricity,heat,renewable,consume by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
48. power
Collect: gather authentic examples containing power. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as energy,control,authority,electricity,strength by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
49. spring
Collect: gather authentic examples containing spring. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as season,flower,warm,summer,winter by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
50. spring
Collect: gather authentic examples containing spring. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as water,source,river,natural,flow by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
51. virus
Collect: gather authentic examples containing virus. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as infection,disease,cell,spread,immune by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
52. virus
Collect: gather authentic examples containing virus. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as computer,malware,file,security,infect by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
53. cloud
Collect: gather authentic examples containing cloud. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as rain,sky,weather,storm,grey by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
54. cloud
Collect: gather authentic examples containing cloud. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as server,data,storage,computing,service by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
55. root
Collect: gather authentic examples containing root. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as plant,soil,tree,grow,stem by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
56. root
Collect: gather authentic examples containing root. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as word,prefix,suffix,morphology,meaning by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
57. field
Collect: gather authentic examples containing field. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as farm,land,crop,grass,farmer by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
58. field
Collect: gather authentic examples containing field. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as research,study,discipline,academic,area by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
59. frame
Collect: gather authentic examples containing frame. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as picture,wood,wall,photo,border by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
60. frame
Collect: gather authentic examples containing frame. Separate examples that belong to different senses before combining them.
Cluster: group recurring neighbours such as scene,role,semantics,background,concept by semantic, grammatical and topical role. The cluster structure is more informative than raw co-occurrence alone.
Compare: choose a near-neighbour and ask which contexts overlap and which separate the two words. This turns distribution into a precision tool rather than a synonym generator.
Produce: write a new sentence matching the typical distribution while preserving the target’s exact meaning and register.
27. Frequently asked questions
What is distributional semantics?
A family of approaches that represent meaning from patterns of linguistic distribution in corpora.
What is the distributional hypothesis?
The idea that words occurring in similar contexts tend to have related semantic properties.
What is co-occurrence?
The tendency of words to appear near or in structured relation with one another.
What is a semantic vector?
A numerical representation summarising contextual distribution.
What is cosine similarity?
A common measure comparing the direction of two vectors in semantic space.
Is distributional similarity the same as synonymy?
No. It may reflect category similarity, topical relatedness, functional similarity or other shared context.
How does polysemy affect models?
One representation can blur several senses unless contexts are separated.
Can distributional semantics fully explain meaning?
No. Textual distribution captures important semantic information but not complete experiential grounding.
Why does corpus choice matter?
Different corpora expose different senses, registers, topics and biases.
What is the simplest rule?
Look at the company a word repeatedly keeps, then verify what that company really tells you.
28. Research grounding
Cambridge’s 2023 book Distributional Semantics describes the field as a theoretical, computational and cognitive framework in which semantic representations are constructed from statistical distribution in linguistic contexts. Cambridge’s review of word and sense similarity likewise states the distributional assumption that semantic properties can be inferred from co-occurrence patterns in corpora. Recent Cambridge work on the neuroscience of word meaning notes that distributional semantic models capture important aspects of conceptual processing while remaining incomplete accounts of grounded meaning.
- Cambridge — Distributional Semantics
- Cambridge — Distributional word and sense similarity
- Cambridge — Distributional models and word meaning
29. eduKateSG routes
30. Final model
Distributional semantics shows that language use itself carries structured evidence about meaning. Repeated contextual patterns reveal categories, relations, senses and expectations.
The strongest vocabulary learner uses that evidence without confusing it for the whole concept: corpus patterns tell us how words behave, while world knowledge tells us what those patterns are about.
Meaning is not identical to distribution—but distribution is one of meaning’s clearest traces.
31. Final distributional audit
1. doctor
Counterfactual corpus: imagine a corpus where doctor appears only in healthcare contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile nurse,hospital,patient,clinic,medicine to build a teaching set, then remove any association that would encourage inaccurate meaning.
2. teacher
Counterfactual corpus: imagine a corpus where teacher appears only in education contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile student,classroom,lesson,school,teach to build a teaching set, then remove any association that would encourage inaccurate meaning.
3. knife
Counterfactual corpus: imagine a corpus where knife appears only in tool/action contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cut,sharp,blade,kitchen,slice to build a teaching set, then remove any association that would encourage inaccurate meaning.
4. scissors
Counterfactual corpus: imagine a corpus where scissors appears only in tool/action contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cut,snip,paper,hair,sharp to build a teaching set, then remove any association that would encourage inaccurate meaning.
5. rain
Counterfactual corpus: imagine a corpus where rain appears only in weather contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cloud,wet,storm,weather,umbrella to build a teaching set, then remove any association that would encourage inaccurate meaning.
6. snow
Counterfactual corpus: imagine a corpus where snow appears only in weather contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cold,winter,ice,white,weather to build a teaching set, then remove any association that would encourage inaccurate meaning.
7. bank
Counterfactual corpus: imagine a corpus where bank appears only in financial sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile money,loan,account,finance,interest to build a teaching set, then remove any association that would encourage inaccurate meaning.
8. bank
Counterfactual corpus: imagine a corpus where bank appears only in river-edge sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile river,shore,water,stream,flood to build a teaching set, then remove any association that would encourage inaccurate meaning.
9. cell
Counterfactual corpus: imagine a corpus where cell appears only in biological sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile biology,membrane,nucleus,tissue,organism to build a teaching set, then remove any association that would encourage inaccurate meaning.
10. cell
Counterfactual corpus: imagine a corpus where cell appears only in prison sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile prison,inmate,jail,locked,guard to build a teaching set, then remove any association that would encourage inaccurate meaning.
11. model
Counterfactual corpus: imagine a corpus where model appears only in computational sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,predict,training,system,parameter to build a teaching set, then remove any association that would encourage inaccurate meaning.
12. model
Counterfactual corpus: imagine a corpus where model appears only in person/fashion sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile fashion,photograph,runway,agency,pose to build a teaching set, then remove any association that would encourage inaccurate meaning.
13. run
Counterfactual corpus: imagine a corpus where run appears only in movement sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile race,fast,walk,exercise,track to build a teaching set, then remove any association that would encourage inaccurate meaning.
14. run
Counterfactual corpus: imagine a corpus where run appears only in management sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile business,manage,company,operate,organisation to build a teaching set, then remove any association that would encourage inaccurate meaning.
15. light
Counterfactual corpus: imagine a corpus where light appears only in illumination sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile bright,dark,lamp,shine,room to build a teaching set, then remove any association that would encourage inaccurate meaning.
16. light
Counterfactual corpus: imagine a corpus where light appears only in weight sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile weight,heavy,carry,thin,portable to build a teaching set, then remove any association that would encourage inaccurate meaning.
17. strong
Counterfactual corpus: imagine a corpus where strong appears only in physical evaluation. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile powerful,weak,strength,muscle,force to build a teaching set, then remove any association that would encourage inaccurate meaning.
18. strong
Counterfactual corpus: imagine a corpus where strong appears only in abstract academic use. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile evidence,argument,case,support,result to build a teaching set, then remove any association that would encourage inaccurate meaning.
19. heavy
Counterfactual corpus: imagine a corpus where heavy appears only in physical sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile weight,load,carry,light to build a teaching set, then remove any association that would encourage inaccurate meaning.
20. heavy
Counterfactual corpus: imagine a corpus where heavy appears only in abstract/collocational use. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rain,traffic,pressure,loss to build a teaching set, then remove any association that would encourage inaccurate meaning.
21. evidence
Counterfactual corpus: imagine a corpus where evidence appears only in academic reasoning. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile support,claim,data,result,proof,study to build a teaching set, then remove any association that would encourage inaccurate meaning.
22. hypothesis
Counterfactual corpus: imagine a corpus where hypothesis appears only in science/research. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile test,theory,predict,experiment,evidence to build a teaching set, then remove any association that would encourage inaccurate meaning.
23. algorithm
Counterfactual corpus: imagine a corpus where algorithm appears only in computing. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,compute,model,search,process to build a teaching set, then remove any association that would encourage inaccurate meaning.
24. enzyme
Counterfactual corpus: imagine a corpus where enzyme appears only in biology. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile protein,reaction,catalyst,cell,substrate to build a teaching set, then remove any association that would encourage inaccurate meaning.
25. inflation
Counterfactual corpus: imagine a corpus where inflation appears only in economics. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile prices,economy,rate,cost,central bank to build a teaching set, then remove any association that would encourage inaccurate meaning.
26. jurisdiction
Counterfactual corpus: imagine a corpus where jurisdiction appears only in law. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile court,law,authority,case,legal to build a teaching set, then remove any association that would encourage inaccurate meaning.
27. tempo
Counterfactual corpus: imagine a corpus where tempo appears only in music. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile music,beat,rhythm,speed,performance to build a teaching set, then remove any association that would encourage inaccurate meaning.
28. offside
Counterfactual corpus: imagine a corpus where offside appears only in sport. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile football,referee,goal,line,player to build a teaching set, then remove any association that would encourage inaccurate meaning.
29. selfie
Counterfactual corpus: imagine a corpus where selfie appears only in digital culture. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile photo,phone,camera,post,social media to build a teaching set, then remove any association that would encourage inaccurate meaning.
30. podcast
Counterfactual corpus: imagine a corpus where podcast appears only in media. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile episode,audio,listen,host,show to build a teaching set, then remove any association that would encourage inaccurate meaning.
31. happy
Counterfactual corpus: imagine a corpus where happy appears only in emotion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile joy,smile,glad,pleased,sad to build a teaching set, then remove any association that would encourage inaccurate meaning.
32. angry
Counterfactual corpus: imagine a corpus where angry appears only in emotion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile mad,furious,annoyed,rage,upset to build a teaching set, then remove any association that would encourage inaccurate meaning.
33. walk
Counterfactual corpus: imagine a corpus where walk appears only in motion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile run,stroll,road,feet,move to build a teaching set, then remove any association that would encourage inaccurate meaning.
34. whisper
Counterfactual corpus: imagine a corpus where whisper appears only in speech manner. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile speak,quiet,voice,soft,talk to build a teaching set, then remove any association that would encourage inaccurate meaning.
35. purchase
Counterfactual corpus: imagine a corpus where purchase appears only in commerce. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile buy,price,customer,sale,payment to build a teaching set, then remove any association that would encourage inaccurate meaning.
36. buy
Counterfactual corpus: imagine a corpus where buy appears only in commerce. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile sell,money,shop,pay,price to build a teaching set, then remove any association that would encourage inaccurate meaning.
37. child
Counterfactual corpus: imagine a corpus where child appears only in human/social. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile parent,school,young,family,play to build a teaching set, then remove any association that would encourage inaccurate meaning.
38. infant
Counterfactual corpus: imagine a corpus where infant appears only in human/social. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile baby,child,birth,mother,newborn to build a teaching set, then remove any association that would encourage inaccurate meaning.
39. vehicle
Counterfactual corpus: imagine a corpus where vehicle appears only in category. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile car,truck,bus,transport,road to build a teaching set, then remove any association that would encourage inaccurate meaning.
40. car
Counterfactual corpus: imagine a corpus where car appears only in category member. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile vehicle,drive,road,wheel,engine to build a teaching set, then remove any association that would encourage inaccurate meaning.
41. fruit
Counterfactual corpus: imagine a corpus where fruit appears only in category. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile apple,banana,orange,food,tree to build a teaching set, then remove any association that would encourage inaccurate meaning.
42. apple
Counterfactual corpus: imagine a corpus where apple appears only in category member. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile fruit,red,tree,eat,juice to build a teaching set, then remove any association that would encourage inaccurate meaning.
43. justice
Counterfactual corpus: imagine a corpus where justice appears only in abstract social concept. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile law,fair,rights,court,equality to build a teaching set, then remove any association that would encourage inaccurate meaning.
44. freedom
Counterfactual corpus: imagine a corpus where freedom appears only in abstract social concept. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rights,liberty,choice,independence,democracy to build a teaching set, then remove any association that would encourage inaccurate meaning.
45. data
Counterfactual corpus: imagine a corpus where data appears only in research/computing. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile analysis,information,collect,result,statistics to build a teaching set, then remove any association that would encourage inaccurate meaning.
46. information
Counterfactual corpus: imagine a corpus where information appears only in general abstract noun. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,knowledge,source,provide,detail to build a teaching set, then remove any association that would encourage inaccurate meaning.
47. energy
Counterfactual corpus: imagine a corpus where energy appears only in science/public. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile power,electricity,heat,renewable,consume to build a teaching set, then remove any association that would encourage inaccurate meaning.
48. power
Counterfactual corpus: imagine a corpus where power appears only in polysemous general term. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile energy,control,authority,electricity,strength to build a teaching set, then remove any association that would encourage inaccurate meaning.
49. spring
Counterfactual corpus: imagine a corpus where spring appears only in season sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile season,flower,warm,summer,winter to build a teaching set, then remove any association that would encourage inaccurate meaning.
50. spring
Counterfactual corpus: imagine a corpus where spring appears only in water-source sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile water,source,river,natural,flow to build a teaching set, then remove any association that would encourage inaccurate meaning.
51. virus
Counterfactual corpus: imagine a corpus where virus appears only in biology. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile infection,disease,cell,spread,immune to build a teaching set, then remove any association that would encourage inaccurate meaning.
52. virus
Counterfactual corpus: imagine a corpus where virus appears only in computing metaphor. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile computer,malware,file,security,infect to build a teaching set, then remove any association that would encourage inaccurate meaning.
53. cloud
Counterfactual corpus: imagine a corpus where cloud appears only in weather sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rain,sky,weather,storm,grey to build a teaching set, then remove any association that would encourage inaccurate meaning.
54. cloud
Counterfactual corpus: imagine a corpus where cloud appears only in computing sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile server,data,storage,computing,service to build a teaching set, then remove any association that would encourage inaccurate meaning.
55. root
Counterfactual corpus: imagine a corpus where root appears only in botanical sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile plant,soil,tree,grow,stem to build a teaching set, then remove any association that would encourage inaccurate meaning.
56. root
Counterfactual corpus: imagine a corpus where root appears only in linguistic sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile word,prefix,suffix,morphology,meaning to build a teaching set, then remove any association that would encourage inaccurate meaning.
57. field
Counterfactual corpus: imagine a corpus where field appears only in physical sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile farm,land,crop,grass,farmer to build a teaching set, then remove any association that would encourage inaccurate meaning.
58. field
Counterfactual corpus: imagine a corpus where field appears only in abstract domain sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile research,study,discipline,academic,area to build a teaching set, then remove any association that would encourage inaccurate meaning.
59. frame
Counterfactual corpus: imagine a corpus where frame appears only in object sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile picture,wood,wall,photo,border to build a teaching set, then remove any association that would encourage inaccurate meaning.
60. frame
Counterfactual corpus: imagine a corpus where frame appears only in linguistic/cognitive sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile scene,role,semantics,background,concept to build a teaching set, then remove any association that would encourage inaccurate meaning.
61. doctor
Counterfactual corpus: imagine a corpus where doctor appears only in healthcare contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile nurse,hospital,patient,clinic,medicine to build a teaching set, then remove any association that would encourage inaccurate meaning.
62. teacher
Counterfactual corpus: imagine a corpus where teacher appears only in education contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile student,classroom,lesson,school,teach to build a teaching set, then remove any association that would encourage inaccurate meaning.
63. knife
Counterfactual corpus: imagine a corpus where knife appears only in tool/action contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cut,sharp,blade,kitchen,slice to build a teaching set, then remove any association that would encourage inaccurate meaning.
64. scissors
Counterfactual corpus: imagine a corpus where scissors appears only in tool/action contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cut,snip,paper,hair,sharp to build a teaching set, then remove any association that would encourage inaccurate meaning.
65. rain
Counterfactual corpus: imagine a corpus where rain appears only in weather contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cloud,wet,storm,weather,umbrella to build a teaching set, then remove any association that would encourage inaccurate meaning.
66. snow
Counterfactual corpus: imagine a corpus where snow appears only in weather contexts. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile cold,winter,ice,white,weather to build a teaching set, then remove any association that would encourage inaccurate meaning.
67. bank
Counterfactual corpus: imagine a corpus where bank appears only in financial sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile money,loan,account,finance,interest to build a teaching set, then remove any association that would encourage inaccurate meaning.
68. bank
Counterfactual corpus: imagine a corpus where bank appears only in river-edge sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile river,shore,water,stream,flood to build a teaching set, then remove any association that would encourage inaccurate meaning.
69. cell
Counterfactual corpus: imagine a corpus where cell appears only in biological sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile biology,membrane,nucleus,tissue,organism to build a teaching set, then remove any association that would encourage inaccurate meaning.
70. cell
Counterfactual corpus: imagine a corpus where cell appears only in prison sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile prison,inmate,jail,locked,guard to build a teaching set, then remove any association that would encourage inaccurate meaning.
71. model
Counterfactual corpus: imagine a corpus where model appears only in computational sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,predict,training,system,parameter to build a teaching set, then remove any association that would encourage inaccurate meaning.
72. model
Counterfactual corpus: imagine a corpus where model appears only in person/fashion sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile fashion,photograph,runway,agency,pose to build a teaching set, then remove any association that would encourage inaccurate meaning.
73. run
Counterfactual corpus: imagine a corpus where run appears only in movement sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile race,fast,walk,exercise,track to build a teaching set, then remove any association that would encourage inaccurate meaning.
74. run
Counterfactual corpus: imagine a corpus where run appears only in management sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile business,manage,company,operate,organisation to build a teaching set, then remove any association that would encourage inaccurate meaning.
75. light
Counterfactual corpus: imagine a corpus where light appears only in illumination sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile bright,dark,lamp,shine,room to build a teaching set, then remove any association that would encourage inaccurate meaning.
76. light
Counterfactual corpus: imagine a corpus where light appears only in weight sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile weight,heavy,carry,thin,portable to build a teaching set, then remove any association that would encourage inaccurate meaning.
77. strong
Counterfactual corpus: imagine a corpus where strong appears only in physical evaluation. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile powerful,weak,strength,muscle,force to build a teaching set, then remove any association that would encourage inaccurate meaning.
78. strong
Counterfactual corpus: imagine a corpus where strong appears only in abstract academic use. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile evidence,argument,case,support,result to build a teaching set, then remove any association that would encourage inaccurate meaning.
79. heavy
Counterfactual corpus: imagine a corpus where heavy appears only in physical sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile weight,load,carry,light to build a teaching set, then remove any association that would encourage inaccurate meaning.
80. heavy
Counterfactual corpus: imagine a corpus where heavy appears only in abstract/collocational use. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rain,traffic,pressure,loss to build a teaching set, then remove any association that would encourage inaccurate meaning.
81. evidence
Counterfactual corpus: imagine a corpus where evidence appears only in academic reasoning. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile support,claim,data,result,proof,study to build a teaching set, then remove any association that would encourage inaccurate meaning.
82. hypothesis
Counterfactual corpus: imagine a corpus where hypothesis appears only in science/research. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile test,theory,predict,experiment,evidence to build a teaching set, then remove any association that would encourage inaccurate meaning.
83. algorithm
Counterfactual corpus: imagine a corpus where algorithm appears only in computing. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,compute,model,search,process to build a teaching set, then remove any association that would encourage inaccurate meaning.
84. enzyme
Counterfactual corpus: imagine a corpus where enzyme appears only in biology. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile protein,reaction,catalyst,cell,substrate to build a teaching set, then remove any association that would encourage inaccurate meaning.
85. inflation
Counterfactual corpus: imagine a corpus where inflation appears only in economics. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile prices,economy,rate,cost,central bank to build a teaching set, then remove any association that would encourage inaccurate meaning.
86. jurisdiction
Counterfactual corpus: imagine a corpus where jurisdiction appears only in law. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile court,law,authority,case,legal to build a teaching set, then remove any association that would encourage inaccurate meaning.
87. tempo
Counterfactual corpus: imagine a corpus where tempo appears only in music. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile music,beat,rhythm,speed,performance to build a teaching set, then remove any association that would encourage inaccurate meaning.
88. offside
Counterfactual corpus: imagine a corpus where offside appears only in sport. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile football,referee,goal,line,player to build a teaching set, then remove any association that would encourage inaccurate meaning.
89. selfie
Counterfactual corpus: imagine a corpus where selfie appears only in digital culture. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile photo,phone,camera,post,social media to build a teaching set, then remove any association that would encourage inaccurate meaning.
90. podcast
Counterfactual corpus: imagine a corpus where podcast appears only in media. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile episode,audio,listen,host,show to build a teaching set, then remove any association that would encourage inaccurate meaning.
91. happy
Counterfactual corpus: imagine a corpus where happy appears only in emotion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile joy,smile,glad,pleased,sad to build a teaching set, then remove any association that would encourage inaccurate meaning.
92. angry
Counterfactual corpus: imagine a corpus where angry appears only in emotion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile mad,furious,annoyed,rage,upset to build a teaching set, then remove any association that would encourage inaccurate meaning.
93. walk
Counterfactual corpus: imagine a corpus where walk appears only in motion. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile run,stroll,road,feet,move to build a teaching set, then remove any association that would encourage inaccurate meaning.
94. whisper
Counterfactual corpus: imagine a corpus where whisper appears only in speech manner. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile speak,quiet,voice,soft,talk to build a teaching set, then remove any association that would encourage inaccurate meaning.
95. purchase
Counterfactual corpus: imagine a corpus where purchase appears only in commerce. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile buy,price,customer,sale,payment to build a teaching set, then remove any association that would encourage inaccurate meaning.
96. buy
Counterfactual corpus: imagine a corpus where buy appears only in commerce. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile sell,money,shop,pay,price to build a teaching set, then remove any association that would encourage inaccurate meaning.
97. child
Counterfactual corpus: imagine a corpus where child appears only in human/social. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile parent,school,young,family,play to build a teaching set, then remove any association that would encourage inaccurate meaning.
98. infant
Counterfactual corpus: imagine a corpus where infant appears only in human/social. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile baby,child,birth,mother,newborn to build a teaching set, then remove any association that would encourage inaccurate meaning.
99. vehicle
Counterfactual corpus: imagine a corpus where vehicle appears only in category. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile car,truck,bus,transport,road to build a teaching set, then remove any association that would encourage inaccurate meaning.
100. car
Counterfactual corpus: imagine a corpus where car appears only in category member. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile vehicle,drive,road,wheel,engine to build a teaching set, then remove any association that would encourage inaccurate meaning.
101. fruit
Counterfactual corpus: imagine a corpus where fruit appears only in category. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile apple,banana,orange,food,tree to build a teaching set, then remove any association that would encourage inaccurate meaning.
102. apple
Counterfactual corpus: imagine a corpus where apple appears only in category member. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile fruit,red,tree,eat,juice to build a teaching set, then remove any association that would encourage inaccurate meaning.
103. justice
Counterfactual corpus: imagine a corpus where justice appears only in abstract social concept. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile law,fair,rights,court,equality to build a teaching set, then remove any association that would encourage inaccurate meaning.
104. freedom
Counterfactual corpus: imagine a corpus where freedom appears only in abstract social concept. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rights,liberty,choice,independence,democracy to build a teaching set, then remove any association that would encourage inaccurate meaning.
105. data
Counterfactual corpus: imagine a corpus where data appears only in research/computing. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile analysis,information,collect,result,statistics to build a teaching set, then remove any association that would encourage inaccurate meaning.
106. information
Counterfactual corpus: imagine a corpus where information appears only in general abstract noun. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile data,knowledge,source,provide,detail to build a teaching set, then remove any association that would encourage inaccurate meaning.
107. energy
Counterfactual corpus: imagine a corpus where energy appears only in science/public. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile power,electricity,heat,renewable,consume to build a teaching set, then remove any association that would encourage inaccurate meaning.
108. power
Counterfactual corpus: imagine a corpus where power appears only in polysemous general term. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile energy,control,authority,electricity,strength to build a teaching set, then remove any association that would encourage inaccurate meaning.
109. spring
Counterfactual corpus: imagine a corpus where spring appears only in season sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile season,flower,warm,summer,winter to build a teaching set, then remove any association that would encourage inaccurate meaning.
110. spring
Counterfactual corpus: imagine a corpus where spring appears only in water-source sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile water,source,river,natural,flow to build a teaching set, then remove any association that would encourage inaccurate meaning.
111. virus
Counterfactual corpus: imagine a corpus where virus appears only in biology. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile infection,disease,cell,spread,immune to build a teaching set, then remove any association that would encourage inaccurate meaning.
112. virus
Counterfactual corpus: imagine a corpus where virus appears only in computing metaphor. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile computer,malware,file,security,infect to build a teaching set, then remove any association that would encourage inaccurate meaning.
113. cloud
Counterfactual corpus: imagine a corpus where cloud appears only in weather sense. Predict which neighbours would become artificially dominant and which senses would disappear.
Human-versus-vector check: identify one relation humans know from world experience but a text-only distributional model may represent weakly. This exposes the boundary between linguistic evidence and grounded concept knowledge.
Vocabulary decision: use the profile rain,sky,weather,storm,grey to build a teaching set, then remove any association that would encourage inaccurate meaning.
