Voynich words have a family problem.
Look at enough lines and a strange sensation develops.
You have seen this word before.
Not exactly.
Almost.
One glyph has changed.
A short beginning has been added.
An ending has disappeared.
A central group remains while the edges change.
Two neighbouring tokens look like siblings.
Then the next line produces another cousin.
This is one of the most famous structural properties of Voynichese.
Its word-like units do not feel freely assembled from a large alphabet.
They feel constrained.
Many forms occupy recurring families.
Small edits turn one common form into another.
Local passages can contain strings that resemble one another more strongly than words randomly selected from distant parts of the manuscript.
This observation has generated radically different explanations.
Perhaps we are seeing morphology.
Root plus prefix.
Root plus suffix.
Case, number, tense, derivation.
Perhaps we are seeing medieval abbreviation.
One underlying word receives several compressed forms.
Perhaps we are seeing ciphertext.
Related visible groups arise from encoding rules rather than related plaintext words.
Perhaps we are seeing a local generation process in which one token is modified to create another.
Torsten Timm has argued strongly for that last family of explanations, proposing that similarly spelled words reveal a text-generation mechanism.
Other researchers have treated the same family structure as compatible with language, structured notation or encoding.
The useful fact comes before all of them.
Voynich tokens live in a highly constrained neighbourhood of possible forms.
The problem is discovering what kind of neighbourhood it is.
Quick Read
One-sentence answer: Voynich word-like units form unusually tight families in which many common tokens can be related through small additions, deletions or substitutions, but this structure can arise from natural-language morphology, abbreviation, positional orthography, encoding rules or constrained generation, so word similarity is a major constraint on explanation rather than a decipherment by itself.
- Voynich visible spaces create convenient token units, but those units are not proven lexical words.
- Many tokens share recurring internal components and differ by short prefixes, suffixes or glyph substitutions.
- Common EVA strings such as forms in the chedy, qokeedy, daiin and related families illustrate the phenomenon without revealing meaning.
- The inventory of possible glyph positions inside a token is highly constrained.
- Some glyphs strongly prefer beginnings, middles or endings of tokens.
- Nearby words can be unusually similar, making local neighbourhood effects important.
- Small edit distance does not automatically mean grammatical inflection.
- Natural languages routinely create families through morphology, derivation, clitics and spelling variation.
- Historical abbreviation systems can also create surface families from related or repeated underlying forms.
- Verbose or homophonic encoding can produce related-looking visible units whose plaintext relationship is not obvious.
- Generation models can intentionally modify one token into nearby variants, naturally producing dense families.
- Line position interacts with families because line beginnings and endings can add or favour particular components.
- Currier A/B interacts with families because some endings and groups are strongly regime-biased.
- Transcription and segmentation matter: splitting one compound changes edit distance and family structure.
- The best explanatory model should predict which variants occur, where they occur and what independent variable changes with them.
“These words are similar” is therefore only the first sentence.
The important sentence is:
what operation turns one family member into another?
First: “Word” Is a Convenience
Voynich writing contains visible spaces.
Researchers naturally treat space-delimited strings as words.
This is useful.
It may even be correct.
But the word “word” carries linguistic assumptions.
A visible group could be:
- a lexical word;
- a syllable group;
- a morpheme bundle;
- a code group;
- a cipher chunk;
- a generated unit;
- a scribal rhythm unit.
So throughout this article, “word family” means:
a family of visually space-delimited Voynich tokens related by recurring glyph structure.
Semantics comes later.
Edit Distance Gives Us a Neutral Way to Talk About “Almost the Same”
Suppose we have two strings:
ABCDE
ABCDF
Only one symbol changes.
The edit distance is small.
Or:
ABCDE
XABCDE
One symbol is inserted.
Again, the strings are close.
Edit distance lets us measure similarity without first claiming that one unit is a prefix or suffix.
This is valuable because grammatical language is only one possible cause of small edits.
The same mathematical relationship can arise from copying, mutation, encoding or generation.
Distance describes.
Mechanism explains.
Voynich Words Occupy Restricted Slots
One reason word families feel so strong is that not every glyph can appear anywhere.
Some forms strongly prefer token beginnings.
Some strongly prefer endings.
Some combinations appear in stable internal positions.
This gives Voynich tokens a slot-like character.
A rough schematic might look like:
[optional opening] + [core family] + [optional ending]
This is not a decoded morphology.
It is a description of positional regularity.
Natural languages can have templatic or affix-heavy morphology.
Ciphers can have slots.
Generated systems can have slots.
The slot structure gives every theory a common object to explain.
Natural Language Creates Word Families All the Time
English gives us:
- teach;
- teacher;
- teaches;
- teaching;
- reteach.
One core participates in several related forms.
Latin and other inflected languages can produce even denser families through case, number, gender and verb endings.
Agglutinative systems can construct long forms from stable sequences of morphemes.
Therefore dense word families are compatible with natural language.
The question is whether Voynich families behave like meaningful morphology.
Do endings correlate with grammatical environments?
Do variants recur in syntactically predictable positions?
Does one family correspond to one semantic root across contexts?
Those are much stronger tests than resemblance alone.
Morphology Predicts Meaningful Relationships
If two forms differ by a plural ending, their meanings should be related.
If two forms differ by tense, their syntactic environments should reflect that.
If a prefix changes meaning systematically, the same prefix should have related effects across many roots.
This gives a natural-language morphology hypothesis predictive content.
A future decipherment should be able to say:
adding this visible ending corresponds to this grammatical operation.
Then the operation should work on forms that were not used to discover it.
That is how a visual family becomes linguistic morphology.
But Voynich Families Can Be Too Local for Simple Lexical Morphology
One of the stranger properties of Voynich vocabulary is locality.
Some token forms cluster strongly on particular pages or nearby regions.
This can be normal if nearby text discusses one topic.
A botanical passage repeats plant-specific terms.
A recipe repeats ingredients.
But extreme local similarity can also arise from local generation or copying processes.
This is why locality is not simply “topic evidence”.
The same observation can be semantic or mechanical.
A good model must explain not merely that words are local, but how local their family relationships are compared with realistic language and encoding controls.
Medieval Abbreviation Can Also Create Families
Medieval manuscripts often compress frequently used words and endings.
A visible mark can stand for several letters.
One root can appear in a full form and several abbreviated forms.
Line position can influence abbreviation choices.
This matters because some Voynich variation could reflect an underlying language whose surface representation is compressed.
An abbreviation model predicts historical regularity.
Variants should expand under stable conventions.
One mark should not stand for whatever missing letters are needed on each page.
Unlimited abbreviation makes every language easy to find.
Constrained abbreviation can become a real decipherment mechanism.
Encoding Can Create Surface Families That Are Not Plaintext Families
Suppose one plaintext letter can be encoded by several visible groups.
Or one plaintext syllable is expanded into a longer visible pattern.
Now many ciphertext groups can share structural components without representing related plaintext words.
A homophonic or verbose system can therefore produce families.
The 2025 Naibbe study is relevant here because it demonstrates that a historically plausible hand-operable encoding can reproduce several Voynich-like statistical properties from meaningful Latin or Italian plaintext.
That does not mean Voynich word families are ciphertext families.
It means the cipher alternative cannot be dismissed merely because the visible forms look internally structured.
An encoding theory should specify how the visible family is generated from underlying units and then predict unseen forms.
Local Generation Is the Most Radical Word-Family Explanation
Torsten Timm’s work argues that similarly spelled Voynich words can be explained by a production method in which existing forms are copied and modified.
Under a simplified version of this idea:
- choose a nearby or available token;
- add, remove or substitute a glyph;
- write the new variant;
- use it as material for another variant.
This naturally creates dense local families.
It also naturally produces many one-edit neighbours.
And it can make vocabulary drift gradually across a manuscript.
The model is attractive because it explains surface structure directly.
Its burden appears elsewhere.
Why does the generated text exhibit long-range organisation?
Why do page families differ systematically?
Why is the text integrated with hundreds of images and labels?
If the process has meaning, how is meaning maintained during local mutation?
If it has no meaning, what was the function of the resulting object?
The generation hypothesis is therefore powerful at one layer and incomplete at another.
Generation Does Not Automatically Mean Hoax
This distinction matters.
A text can be generated mechanically and still serve a function.
A mnemonic system can generate constrained labels.
A coded notation can generate formulaic strings.
A ritual or combinatorial system can create patterned text meaningful to trained users.
Therefore evidence for local generation would not by itself establish meaningless gibberish.
We would still need to ask:
what does the generation procedure preserve?
Meaning?
Class?
Quantity?
Indexing?
Nothing semantic?
“Generated” is a production description, not a complete functional interpretation.
Copying From an Exemplar Can Mimic Local Generation
There is another possibility between language and free generation.
A scribe may copy from a source whose entries are already related.
A dictionary has alphabetic neighbours.
A medical list groups related substances.
A paradigm lists grammatical variants together.
A codebook groups related codes.
Local word similarity can therefore reflect source organisation rather than on-the-fly mutation.
This matters because manuscript order itself may be inherited from an exemplar.
The visible local family is compatible with multiple production stories.
Source structure is an underappreciated alternative.
Line Position Can Manufacture Family Members
Now connect this article to the line article.
Suppose an ordinary internal form becomes an expanded form at line start.
Or receives a special ending at line end.
Those positional variants now look like members of one word family.
Are they grammatical derivatives?
Or one underlying unit written differently because of physical position?
Without controlling for line position, morphology and scribal layout can be confused.
A strong family analysis therefore asks:
does this edit occur because the meaning changed, or because the token moved to a new structural position?
Currier A/B Can Manufacture Different Family Geometries
Some suffix-like and group-like forms are strongly associated with Currier B, while others characterise A.
This changes which family members are common in each regime.
Imagine one core form.
A uses endings X and Y.
B uses Y and Z.
The same core now sits inside two different visible family geometries.
This could be:
- dialect morphology;
- orthographic variation;
- different encoding state;
- different scribe habit;
- different source vocabulary.
Therefore word-family analysis and Currier analysis cannot be isolated.
A real explanation should tell us why the family rules themselves change across regimes.
Transcription Can Create or Destroy One-Edit Neighbours
Suppose one transcriber treats a connected form as one glyph.
Another splits it into two.
Two tokens that differ by one edit under the first representation may differ by two under the second.
Likewise, an uncertain space can split one token into two and completely change family counts.
This is why edit-distance claims should specify the transcription and segmentation used.
The strongest family phenomena should survive reasonable representations.
If a result disappears under another major transliteration, it may be partly an artefact of encoding the manuscript into data.
Token Similarity Is Not the Same as Semantic Similarity
English gives us cat and car.
One letter differs.
The meanings are not closely related.
It also gives us teach and teacher.
Small edit.
Strong semantic relation.
Surface similarity alone cannot distinguish these cases.
Voynich has the same problem at larger scale.
A one-edit neighbour may be:
- grammatical relative;
- unrelated lexical item;
- encoded variant;
- copying variant;
- generated mutation.
Meaning must be established independently before edit distance can become morphology.
The Nearest Neighbour Is Not Always the Parent
Generation models often imagine a local token being modified into another.
But similarity does not establish direction.
If A and B differ by one edit, we do not know whether:
- A generated B;
- B generated A;
- both derive from C;
- they are independent members of a morphological paradigm;
- they only happen to be similar.
Direction requires chronology or a production rule.
Current reading order cannot safely provide that by itself because page order is not guaranteed production order.
A similarity graph is not a genealogy until direction is independently justified.
Local Repetition Can Be Semantic
Do not forget the ordinary explanation.
Meaningful texts repeat locally.
A paragraph about blood repeats blood.
A plant entry repeats plant names or properties.
A recipe repeats ingredients and operations.
Long-range word-distribution studies have found evidence compatible with topical organisation in Voynich.
So local families may partly reflect content.
The important question is whether the degree and pattern of similarity are plausible under comparable meaningful texts.
Natural-language controls need to preserve genre, length, morphology and historical orthography reasonably well.
A bad control can make Voynich look more exotic than it is.
The Best Model May Be Hybrid
Perhaps the visible families are produced by several layers at once.
An underlying natural language has morphology.
A specialised notation abbreviates it.
A scribe changes forms at line boundaries.
An encoding layer introduces additional variants.
The surface then becomes far more family-dense than ordinary plaintext.
This possibility is realistic.
It is also dangerous because “hybrid” can explain anything.
Each proposed layer must buy specific predictive power.
Otherwise complexity becomes an escape hatch.
Family Structure Should Predict Syntax if It Is Morphology
This is one of the cleanest future tests.
Suppose ending X is proposed as a grammatical suffix.
Then X-bearing words should prefer syntactic environments appropriate to that grammatical category.
If X means plural, neighbouring agreement patterns may change.
If X means a case ending, positional or relational behaviour should reflect it.
If X is simply a line-final variant, its strongest predictor should be margin position instead.
If X is an encoded state, its distribution should follow the cipher rule.
The same visible suffix can therefore make different predictions under different mechanisms.
This is how morphology becomes testable before translation is complete.
Family Structure Should Predict Meaning if It Is Lexical
If several tokens share a semantic root, their contexts should overlap meaningfully.
Suppose a family appears mostly near plant pages with a particular visual class.
That may support a lexical or topical interpretation.
If the same family appears indiscriminately in zodiac, Quire 13, pharmaceutical and recipe contexts, perhaps it is grammatical rather than lexical.
This creates another conditional question:
does family membership predict document context?
Again, a cluster is not a meaning.
But a cluster can generate a test for meaning.
Family Structure Should Predict Production if It Is Generated Locally
A local-generation theory predicts spatial relationships.
Near-neighbour words should be unusually close in form.
Similarity should decay with distance under some models.
Lines may reset local families.
Vocabulary may drift gradually as mutations accumulate.
These are powerful predictions because they concern geometry and sequence rather than meaning.
If they survive across reasonable transcription systems, generation models strengthen.
If family similarity is better predicted by semantic visual section than by local distance, a topic or lexical model strengthens.
Again, the manuscript can be made to choose between mechanisms.
Family Structure Should Predict Encoding if It Is Ciphertext
An encoding theory should be able to generate the visible families from known or plausible plaintext.
This is stronger than saying a complex cipher could produce similar-looking words.
Take a historically plausible mechanism.
Encrypt ordinary period-appropriate text.
Do the resulting token families show comparable:
- edit-distance neighbourhoods;
- prefix and suffix slots;
- line-position behaviour;
- Currier-like state variation;
- conditional entropy?
The Naibbe work is important because it demonstrates that this kind of generative comparison is possible.
A cipher hypothesis becomes much stronger when it reproduces several Voynich family constraints simultaneously rather than one selected statistic.
The Wrong Question Is “Which Theory Can Produce Word Families?”
Many can.
Natural language can.
Abbreviation can.
Ciphertext can.
Local generation can.
The useful question is:
which mechanism predicts the exact geometry of the families, their localness, their line sensitivity, their Currier differences and their relationship to document context?
The details discriminate.
The broad phenomenon does not.
What Survives the Word-Family Work
- Voynich tokens are highly constrained internally.
- Many common forms belong to dense visible families related by small edits.
- Glyph positions within tokens are strongly non-random.
- Local context often contains related forms.
- Line position changes some family variants.
- Currier regime changes the relative frequency of family components.
- Transcription choices can affect family geometry but do not erase the broad phenomenon.
- Natural-language morphology remains viable.
- Historical abbreviation remains viable.
- Ciphertext remains viable.
- Constrained local generation remains viable.
- No one mechanism is established by word similarity alone.
The family structure is one of the strongest constraints on every proposed solution precisely because so many explanations can reach it by different routes.
What Does Not Survive as Established Knowledge
- A proven root-and-affix grammar.
- A decoded prefix inventory.
- A decoded suffix inventory.
- A proof that one-edit neighbours are grammatical variants.
- A proof that local similarity means meaningless generation.
- A proof that generation means hoax.
- A proof that all related-looking tokens share meaning.
- A proof that one token directly generated the next.
- A proof that one family has one modern lexical translation.
The observable family stays.
The semantic dictionary does not.
A Better Word-Family Analysis
- Specify the transliteration and segmentation.
- Define similarity neutrally using edits or shared components.
- Control for token length.
- Control for line position.
- Control for Currier regime.
- Control for page role and visual section.
- Measure local versus distant similarity.
- Separate morphological, abbreviation, encoding and generation predictions.
- Require unseen predictions before assigning semantics.
This keeps the family phenomenon attached to the manuscript instead of attaching it prematurely to one theory.
What Would Count as a Real Word-Family Breakthrough?
Imagine a decipherment establishes that a recurring visible core maps to one stable lexical stem.
Three endings produce three predictable grammatical roles.
The same endings work across dozens of unrelated stems.
The grammatical prediction succeeds on withheld pages.
That would turn visual families into morphology.
Or imagine a historically plausible encoding procedure takes ordinary plaintext and reliably produces the exact same family geometry, line-position variants and Currier-state differences.
That would turn the families into cryptographic evidence.
Or imagine local distance predicts one-edit relationships far better than semantic page category, under several independent transcriptions.
That would strengthen local-generation models.
The breakthrough is the same in every case:
one mechanism begins predicting which family member should appear next and why.
Primary School: Families Without Meaning
Give a child these invented strings:
- mato;
- matos;
- remato;
- maton;
- laki.
Ask which four look related.
The child can identify a family without knowing what any string means.
Then ask:
Does looking related prove the meanings are related?
No.
That is the foundational Voynich lesson.
Lower Secondary: One Edit, Four Causes
Give students two strings that differ by one symbol.
Ask them to invent four mechanisms:
- plural ending;
- copying error;
- cipher variation;
- generated mutation.
All four can produce the same visible relationship.
Then ask what additional evidence would distinguish them.
The student learns that similarity is evidence of relationship structure, not necessarily relationship meaning.
Upper Secondary: Build a Similarity Graph
Give students twenty invented tokens.
Connect two tokens if they differ by one edit.
A graph appears.
Some tokens are hubs.
Some form branches.
Now ask:
Does this graph tell you which token came first?
No.
It gives similarity topology without generation direction.
This is exactly why a Voynich family graph is not automatically a copying genealogy.
JC and Adult Readers: Think in Generative Equivalence
At a higher level, several hidden mechanisms can generate nearly the same surface distribution.
Morphology can create families.
Encoding can create families.
Local mutation can create families.
The inference problem is therefore not merely clustering.
It is finding a statistic or external relation on which the hidden mechanisms are not generatively equivalent.
Syntax may separate morphology from generation.
Spatial decay may separate local generation from lexical morphology.
Historically reproducible ciphertext may test encoding.
The discriminator matters more than the family plot.
A Parent and Teacher Guide
- Notice visible families.
- Describe the edit before assigning grammar.
- Ask whether meaning, position, writer or locality predicts the variant.
- Remember that transcription changes family geometry.
- Do not assume nearest neighbour means parent.
- Compare several generative mechanisms.
- Require a family rule to predict unseen forms before calling it morphology or code.
This teaches a wider habit:
similar outputs do not guarantee similar causes.
Reader Checklist: Before You Interpret a Voynich Word Family
- Are the units true words or only space-delimited tokens?
- Which transcription and segmentation define them?
- How small is the edit distance?
- Could line position explain the difference?
- Could Currier regime explain the difference?
- Could scribal hand explain the difference?
- Is the family local to one page or widespread?
- Does the proposed affix have a consistent effect across many cores?
- Does semantic context support a lexical family?
- Could a plausible encoding produce the same family?
- Could local generation produce it more naturally?
- What observation would distinguish the competing mechanisms?
Frequently Asked Questions
Why do so many Voynich words look alike?
The writing system strongly constrains which glyphs can occupy which positions, producing dense visible families of related token forms. The mechanism behind those families remains debated.
Are these grammatical inflections?
Possibly. Natural-language morphology is a viable explanation, but no accepted mapping has shown that specific visible additions consistently correspond to grammatical operations.
Do similar tokens mean similar things?
Not necessarily. Surface similarity can arise from grammar, spelling, encoding, copying or generation. Semantic similarity requires independent evidence.
What is Torsten Timm’s generation hypothesis?
In broad terms, it proposes that Voynich text can be produced by copying and modifying similar existing tokens, naturally creating local chains of closely related forms. It remains one debated explanation rather than an accepted solution.
Does generated text mean the manuscript is meaningless?
No. A generated or combinatorial notation could still encode categories, references or other structured information. Production method and semantic function are separate questions.
Could the families be ciphertext?
Yes. Historically plausible encoded systems can produce structured visible families. A cipher model would need to reproduce the detailed family, positional and statistical constraints and recover meaningful underlying text.
Why does transcription matter?
Because changing whether a connected form is one glyph or two, or changing a space boundary, changes edit distances and therefore which tokens count as family neighbours.
What is the strongest conclusion?
Voynich tokens occupy a highly constrained structural space with dense families of near-related forms. Any successful explanation of the writing system must account for that family geometry and predict how variants are generated.
Related eduKateSG Reading
- Voynich | Everything eduKate Knows and Tested | The Line as a Unit
- Voynich | Everything eduKate Knows and Tested | Currier A and B
- Voynich | Everything eduKate Knows and Tested | EVA, Transcription and the Segmentation Problem
- Voynich | Everything eduKate Knows and Tested | What the Writing Does Before We Know What It Says
Research and Further Reading
- Torsten Timm — How the Voynich Manuscript Was Created
- Claire Bowern & Luke Lindemann — The Linguistics of the Voynich Manuscript
- Montemurro & Zanette — Keywords and Co-Occurrence Patterns in the Voynich Manuscript
- Michael Greshko — The Naibbe Cipher
- René Zandbergen — Voynich Character and Positional Analysis
The Final Idea
Voynich words keep looking almost alike because the writing system gives them very little freedom.
That is the important fact.
What kind of constraint created the family?
Grammar?
Abbreviation?
Encoding?
Local mutation?
Source organisation?
Several layers at once?
A family tree drawn from visual similarity will not answer that question because similarity has no automatic direction and no automatic semantics.
The answer will arrive when one proposed operation starts predicting the next family member before we see it.
The mystery is not that Voynich words resemble one another. The mystery is what rule keeps them inside the same neighbourhood.
Continue Through the Voynich Research Map
This article is one specialist node in eduKateSG’s larger Voynich research library. Return to the canonical master to see the physical, visual, textual and historical evidence in one continuous argument.
Structure is evidence; structure is not translation. The master keeps resemblance, statistical structure, historical possibility and decipherment claims at separate evidentiary levels.