A label is a tiny piece of text with an unfair advantage.
It sits beside something.
A star.
A human figure.
A zodiac emblem.
A detached plant fragment.
A vessel.
The eye immediately wants to make a dictionary.
This word must name that picture.
If only Voynich research were that easy.
A label can be a name.
It can also be a category.
A number.
A status.
A cross-reference.
An abbreviation.
A code.
A short instruction.
The interesting thing about Voynich labels is that they do not merely look isolated.
As a population, many of them behave differently from ordinary running text.
The most common prose tokens often disappear.
Label vocabulary is unusually flat.
Many labels are unique or nearly unique.
Zodiac and pharmaceutical labels strongly favour some initial character pairs that are much less common in running text.
That is not a translation.
It is something almost as valuable.
The manuscript itself appears to know that labels are a different textual job.
If that result holds under careful representation, every theory of Voynichese has to explain two registers at once.
The prose-like text.
And the label-like text.
Quick Read
One-sentence answer: Voynich labels—especially the large zodiac and pharmaceutical label populations—have vocabulary and frequency properties that differ substantially from ordinary running text, including much flatter frequency distributions and different preferred token openings, suggesting that labels form a purposeful specialised textual register even though their exact functions and meanings remain undeciphered.
- Labels are isolated text units associated spatially with visual objects or diagram components.
- They should be analysed separately from paragraph prose rather than pooled automatically.
- René Zandbergen’s label study found zodiac and pharmaceutical labels have very flat frequency distributions compared with running text.
- In samples reported there, a very high proportion of label tokens are hapax-like or low-frequency forms.
- Several of the most frequent Voynich running-text words are absent or nearly absent as labels.
- Zodiac and pharmaceutical labels strongly favour initial EVA pairs such as ok and ot compared with many running-text populations.
- Running text often favours initial families that labels use much less.
- Zodiac and pharmaceutical label populations resemble each other in important frequency properties despite occurring in different visual sections.
- This makes “labelness” a candidate explanatory variable independent of page topic.
- The result is compatible with labels functioning as names, identifiers, categories or codes.
- It does not prove labels are proper names.
- It does not prove every label describes the adjacent picture literally.
- A meaningless-generation model would need dedicated rules to produce the label register intentionally rather than treating all text as one undifferentiated stream.
- A natural-language, notation or encoding model should predict why labels have a different lexical and structural distribution from prose.
The useful question is not:
What does this one label mean?
It is first:
What kind of textual population do labels form?
What Counts as a Label?
In a manuscript, location can define text type.
A string placed beside one visual object and isolated from prose behaves differently on the page from a string embedded inside a paragraph.
Modern Voynich transliteration formats therefore preserve locus types so that labels can be distinguished from running paragraphs, circular writing and other text geometries.
This matters because once text is flattened into one giant character stream, the distinction disappears.
A label is not a tiny paragraph merely because it contains the same script.
Its document role is different before its meaning is known.
Why Label Statistics Are So Valuable
Decipherment usually struggles because meaning is unknown.
Labels give us a partial functional prior.
We do not know what a label says.
But we know it is acting like a label physically.
That lets us ask whether label text behaves the way specialised names or identifiers often behave.
Do common grammatical-looking words disappear?
Does vocabulary become more unique?
Do particular initial patterns dominate?
Do labels across unrelated visual sections share one register?
These questions can be answered without translation.
That makes labels one of the best places to test whether layout and language interact systematically.
The Frequency Distribution Is Much Flatter
Ordinary running text usually contains a few very common tokens.
In a natural language these are often function words or grammatical forms.
Voynich running text also has frequent recurring tokens.
Label populations behave differently.
Zandbergen’s quantitative comparisons found label distributions unusually flat: relatively few labels recur often, while many occur once or only a few times.
That is exactly the broad pattern we might expect if many labels identify unique entities.
But “might expect” is not “proves”.
Unique codes also have flat distributions.
Catalogue identifiers do.
Number-like labels can.
The flatness tells us the population is unlike ordinary prose.
It does not yet tell us which specialised function is responsible.
Hapax-Heavy Does Not Automatically Mean Proper Names
A hapax is a form occurring once in a sample or corpus.
Proper-name lists often contain many hapaxes because each person or place may have a unique name.
That makes proper names an attractive explanation for Voynich labels.
But other systems generate one-off identifiers too.
- catalogue codes;
- coordinates;
- serial numbers;
- unique abbreviations;
- constructed identifiers.
So the inference is one step at a time.
Flat frequency → specialised register.
Specialised register → names remain plausible.
Names → not yet established.
The Common Prose Words Mostly Disappear
This may be even more informative than label uniqueness.
Some of the most frequent word-like forms in Voynich running text are not used as labels.
If those common forms behave like grammatical glue, their absence from labels makes sense.
A one-word label does not need articles, conjunctions or repeated verbal formulae.
This is compatible with natural-language naming.
It is also compatible with a code system in which running text and labels use different fields or symbol inventories.
The important result is register separation.
The manuscript does not simply take random prose words and place them beside pictures.
Label selection is strongly non-random.
Label Openings Are Different Too
The difference is not only which whole tokens recur.
It appears inside token structure.
Zodiac and pharmaceutical labels favour certain initial EVA pairs, especially families beginning ok and ot, much more strongly than many running-text samples.
Meanwhile, common running-text initial families such as qo, ch and others are much less represented in labels.
This means labels are not merely rare words sampled from the same distribution.
Their internal construction preferences differ.
That is a deeper form of register.
A theory must explain not only why labels choose different vocabulary, but why their word shapes are built differently.
Zodiac and Pharmaceutical Labels Resemble Each Other in the Right Way
This is one of the most important controls.
Zodiac labels occur in one visual world.
Pharmaceutical labels occur beside plant fragments and vessels in another.
If their statistical properties remain more label-like than section-like, document role begins to look causally important.
In Zandbergen’s comparisons, both label populations show unusually flat frequency distributions, with pharmaceutical labels at least as flat as zodiac labels in comparable samples.
That cross-section similarity is valuable because it weakens the simplest topic-only explanation.
Two different picture families may share one textual register because both are being labelled.
This is exactly the kind of crossing evidence that can reveal hidden document structure.
Document Role Can Be More Important Than Topic
Suppose a zodiac label and a pharmaceutical label behave more similarly to each other than either behaves to nearby prose.
Then the hidden variable may be:
label versus running text
rather than:
zodiac versus pharmaceutical topic.
This is a profound correction for section-based analysis.
The manuscript can vary along several axes simultaneously.
- topic;
- scribe;
- Currier regime;
- document role;
- physical sheet;
A cluster may be driven by one or several of them.
Labels give us one of the clearest examples where document role itself has measurable textual consequences.
If Labels Are Names, What Should Follow?
A name hypothesis can be made predictive.
If labels are names of depicted entities:
- the same entity should tend to receive the same or systematically related label;
- related entity classes may share naming morphology;
- common grammatical glue should be rare;
- uniqueness should be high;
- some labels may correspond to historically plausible plant, star, month or personal-name systems.
The first two predictions are especially important.
A dictionary is not built from one picture and one label.
It is built when the naming rule repeats.
If one label is claimed to mean “Mars”, another related label should behave under the same phonology or naming system.
One spectacular match is weak.
A naming system is strong.
If Labels Are Codes, What Should Follow?
A code hypothesis makes different predictions.
Labels may be unique because each identifies one object.
But internal form may reflect category and identifier separately.
For example:
[class prefix] + [unique item code]
This could naturally produce flat vocabulary plus strong recurring initial pairs.
Now ok or ot-like beginnings might mark classes rather than sounds.
Again, that is a hypothetical mechanism.
It becomes evidence only if visual object classes correlate with the proposed prefix classes on unseen labels.
A code theory should predict the pictures.
If Labels Are Categories, Repetition Should Be Selective
Perhaps labels do not name individual entities.
They classify them.
Then repeated labels may appear beside objects sharing one hidden category.
This is easier to test than literal naming in some environments.
Suppose three pharmaceutical plant fragments carry the same label.
Do they share a visual property?
Same root type?
Same vessel association?
Same location on the page?
If repeated labels predict repeated visual classes, category coding strengthens.
If visual objects are unrelated, another function may be more plausible.
If Labels Are Abbreviations, Flat Vocabulary Is Still Possible
A unique underlying name can be heavily abbreviated.
This produces short, unique labels.
It may also produce strong initial patterns if one abbreviation convention is common.
But abbreviation has the same constraint as elsewhere in the manuscript.
The expansion must be systematic.
A label cannot become whichever plant, star or body-part name the adjacent picture suggests.
A successful abbreviation model should predict several labels without using the pictures to choose the expansion first.
A Label Can Be Meaningful Without Naming the Picture
This is one of the most important semantic safeguards.
A diagram label may indicate:
- direction;
- stage;
- material;
- owner;
- number;
- relationship;
- cross-reference;
- action.
Therefore even a perfectly meaningful label need not be the noun for the adjacent object.
This is particularly important on the zodiac pages.
A short token beside a human figure could be:
a name.
a star designation.
a date.
a category.
a house or degree.
Image adjacency creates a relation.
It does not tell us the type of relation.
Labels Are Dangerous Decipherment Anchors Because the Picture Suggests Too Much
Suppose a picture resembles a ram.
A nearby label is short.
The solver searches medieval languages for a word meaning ram or Aries that can be made to match.
After enough spelling flexibility, one appears.
The picture now “confirms” the translation.
But the picture created the candidate.
The validation is circular.
A stronger procedure derives a mapping from other text, freezes it, then applies it to an unseen label.
If the output predicts the visual class before inspection, the image can finally provide independent support.
Labels are excellent validation targets.
They are poor dictionaries when the picture supplies the answer first.
Label Register Is Evidence Against One Undifferentiated Random Generator
If the manuscript text were produced by one simple arbitrary process applied everywhere, why should labels have a different frequency and vocabulary profile from running text?
A meaningless-generation theory can still explain the difference.
It can include a special label-generation mode.
But now the generator must know the document role.
When writing beside an object, it switches vocabulary or formation rules.
That is possible.
It is no longer arbitrary.
The generation mechanism has acquired purposeful state.
This is why label behaviour increases the burden on simple hoax or random-generation models without logically eliminating all generated-text possibilities.
Label Register Is Compatible With Meaningful Language
Natural-language documents routinely have role-specific vocabularies.
Map labels differ from narrative prose.
Figure captions differ from paragraphs.
Personal names differ from sentences.
Ingredient lists differ from explanations.
Therefore the Voynich label distinction is fully compatible with meaningful text.
But compatibility is still not proof.
A meaningful-text theory should eventually recover what role-specific grammar or vocabulary the labels use.
The statistical difference becomes a prediction target for decipherment.
Label Register Is Compatible With Encoding Too
An encoding system can use different modes for names and prose.
Proper names may be encoded with one table.
Common prose with another.
Labels may omit spaces, use codebook entries or rely on shorter groups.
This creates specialised distributions naturally.
But again, a cipher hypothesis must specify the mode switch and reproduce the observed label profile from plausible plaintext labels.
“Cipher could do it” is not enough.
The actual cipher should do it.
Could Labels Reveal the True Word Boundaries?
Because many labels are isolated, they can help test the space problem.
If an isolated label matches a token family appearing inside running text, perhaps the visible prose boundaries are meaningful.
If label forms consistently correspond to subsequences inside longer running tokens, perhaps prose spacing operates at another level.
This creates a useful cross-register comparison.
Labels can serve as naturally isolated candidate units.
They may help reveal whether the same underlying units occur embedded in prose.
Again, the analysis has to control for transcription and segmentation.
Could Labels Reveal Morphology?
If labels are mostly names or nouns, their morphological profile should differ from full prose.
Perhaps some endings common in running text disappear because they encode verbs or grammatical relations.
Perhaps one set of endings concentrates in labels because it marks nominal forms.
This is exactly the kind of distributional evidence that can reveal grammatical categories before vocabulary is translated.
But the inference requires more than “labels are different”.
A proposed nominal ending should recur across many label families and occupy predictable contexts in running text.
If that happens, label register could become a bridge from document role to grammar.
Could Labels Be One of the Best Held-Out Tests?
Yes.
Imagine a decipherment is developed entirely on running text.
The solver freezes the rules.
Then applies them to labels that were not used in model construction.
If the labels resolve into name-like or category-like outputs that independently match the visual objects, this would be powerful.
It would test:
- character mapping;
- word segmentation;
- morphology;
- semantic transfer across registers;
- image-text agreement.
One small label can therefore become a strong validator—but only after it stops being used to invent the answer.
What Survives the Label Work
- Voynich labels form a distinct physical text role.
- Large zodiac and pharmaceutical label populations have markedly flatter frequency distributions than ordinary running text.
- Many common prose tokens are absent from labels.
- Label word-initial patterns differ from many running-text populations.
- Zodiac and pharmaceutical labels share important register-like properties despite different visual domains.
- Document role therefore appears to be an important explanatory variable.
- The observed label behaviour is purposeful enough that any generative model must reproduce it intentionally.
- The data are compatible with meaningful names, identifiers, categories, codes or specialised notation.
- No accepted solution identifies which of those functions dominates.
- Labels are powerful future held-out tests because their visual referents can provide independent validation if they were not used to invent the translation.
What Does Not Survive as Established Knowledge
- Every label is a proper name.
- Every label literally names the adjacent picture.
- ok or ot is a decoded naming prefix.
- Labels prove one specific natural language.
- Labels prove the manuscript is meaningful in the ordinary prose sense.
- Labels rule out all generated-text models.
- Zodiac labels are proven star names or personal names.
- Pharmaceutical labels are proven plant names or ingredients.
The specialised register survives.
The dictionary does not.
A Better Label Analysis
- Define label loci independently of the text itself.
- Keep zodiac, pharmaceutical and other label groups separable.
- Compare frequency distributions with matched running-text samples.
- Compare token-initial and token-final structure.
- Control for Currier regime and hand where relevant.
- Test repeated labels against repeated visual classes.
- Do not use the picture to invent a translation and then call the picture confirmation.
- Use labels as held-out validation targets for models built elsewhere.
- Require any meaningless-generation model to reproduce the specialised label register deliberately.
What Would Count as a Real Label Breakthrough?
Imagine a textual model is built from non-label prose.
It predicts that one common label prefix marks a nominal category.
On unseen zodiac and pharmaceutical labels, the same prefix consistently appears beside one independently defined visual class.
The remaining portions of the labels resolve into stable names or identifiers under the same rules.
That would be a major breakthrough.
Or imagine repeated labels across distant sections identify repeated substances, objects or categories that were not previously recognised as related.
The text would then predict a new visual cross-reference.
That is the kind of cross-modal receipt that can move a label from resemblance to decipherment.
Primary School: A Label Is Not Always a Name
Show a child a classroom diagram.
One arrow says “hot”.
Another says “2”.
Another says “water”.
All are labels.
Only one is the name of a thing.
This immediately breaks the assumption that every Voynich label must name the adjacent object.
Lower Secondary: Compare Two Registers
Take a textbook page.
Make one list of words in the prose paragraph.
Make another list of figure labels.
Compare them.
The prose repeats function words.
The labels contain more nouns and unique terms.
Same language.
Different register.
This gives students an intuitive model for why Voynich label statistics can differ without implying a different language.
Upper Secondary: Topic or Document Role?
Give students four text groups:
- biology prose;
- biology labels;
- astronomy prose;
- astronomy labels.
Ask which groups should resemble each other if topic dominates.
Then ask which should resemble each other if document role dominates.
This is exactly the inference created by zodiac and pharmaceutical label comparisons.
The student learns how crossed categories reveal hidden causes.
JC and Adult Readers: Register Is a Latent Variable
At a higher level, “label” is a document-function variable.
It can alter lexical frequency, morphology and information density independently of language identity.
A model that ignores register may incorrectly attribute label-versus-prose differences to topic, scribe or language.
Voynich therefore requires multi-factor analysis.
Page type is not enough.
Currier is not enough.
Hand is not enough.
Document role can cut across all three.
The more variables are separated before interpretation, the less likely one cluster is to be mistaken for one semantic truth.
A Parent and Teacher Guide
- Separate labels from paragraphs.
- Compare frequency before guessing meaning.
- Notice whether common prose words disappear.
- Ask whether uniqueness suggests names, IDs or categories.
- Compare labels across different topics.
- Do not assume adjacency means naming.
- Use pictures as independent tests, not translation prompts.
The broader lesson is one of the best in this series:
the same writing system can behave differently because the text is doing a different job.
Reader Checklist: Before You Translate a Voynich Label
- Is the string physically a label?
- Which label population does it belong to?
- Does its token shape resemble common label families?
- Is the label unique or repeated?
- Do repeated instances share a visual class?
- Are you assuming it names the adjacent object?
- Did the picture supply the proposed translation?
- Does the same mapping work on other unseen labels?
- Does label morphology differ systematically from prose?
- Could the label be a code, category or number rather than a name?
- Does a proposed generation model reproduce the label register?
- What prediction does the label theory make before the image is inspected?
Frequently Asked Questions
Do Voynich labels behave differently from running text?
Yes. Quantitative comparisons of major label populations show flatter frequency distributions and distinctive vocabulary and initial-character patterns compared with ordinary running text.
Are Voynich labels proper names?
Possibly in some cases, but not proven. Their uniqueness and specialised distribution are compatible with names, identifiers, categories and other label functions.
What is unusual about label frequency?
Many labels are unique or rare, so the distribution is much flatter than prose where a small set of common word-like forms recurs frequently.
Do common Voynich words appear as labels?
Several of the most common running-text tokens are absent or nearly absent from major label populations, which is one reason labels look like a specialised register.
What are the ok/ot patterns?
Zodiac and pharmaceutical labels strongly favour some initial EVA character pairs such as ok and ot compared with many running-text samples. Their function is not decoded.
Do labels prove the text is meaningful?
They strengthen the case that the text-generation process is purposeful and sensitive to document role. They do not by themselves prove ordinary natural-language meaning.
Why are labels useful for decipherment?
Because they create a distinct held-out text population associated with visual objects. A mapping developed elsewhere can be tested on them without using the same picture to invent the answer.
What is the strongest current conclusion?
Labels are a purposefully differentiated textual register in the Voynich Manuscript. Their exact semantic role remains unresolved.
Related eduKateSG Reading
- Voynich | Everything eduKate Knows and Tested | Spaces and Word Boundaries
- Voynich | Everything eduKate Knows and Tested | The Zodiac Pages
- Voynich | Everything eduKate Knows and Tested | The Vessels and Plant Fragments
- Voynich | Everything eduKate Knows and Tested | What a Real Voynich Decipherment Must Survive
Research and Further Reading
- René Zandbergen — Some Special Properties of Labels in the Voynich MS
- René Zandbergen — Transliteration Locus Types and Text Roles
- Claire Bowern & Luke Lindemann — The Linguistics of the Voynich Manuscript
- René Zandbergen — Voynich Transliteration Resources
The Final Idea
Labels are small.
The evidence they carry is not.
They tell us the manuscript does not use its writing system identically everywhere.
The same script changes behaviour when the text changes job.
Prose repeats common forms.
Labels flatten the vocabulary.
Different beginnings become preferred.
Two visually different sections can converge because both are labelling.
That is a powerful structural fact.
It does not tell us whether a label says “Mars”, “red root”, “class three” or “entry 27”.
But it tells us something the final decipherment will have to respect:
The Voynich Manuscript appears to distinguish what text is doing, not merely where text is written.
Continue Through the Voynich Research Map
This article is one specialist node in eduKateSG’s larger Voynich research library. Return to the canonical master to see the physical, visual, textual and historical evidence in one continuous argument.
Structure is evidence; structure is not translation. The master keeps resemblance, statistical structure, historical possibility and decipherment claims at separate evidentiary levels.