VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | Frequency Laws: Zipf, Vocabulary Growth and the “Looks Like Language” Trap

There is a graph that can make an undeciphered manuscript look reassuringly human.

Count every different word-like token.

Rank the tokens from most common to least common.

Plot frequency against rank.

Natural language tends to produce a familiar shape.

A few words are extremely common.

More are moderately common.

A long tail appears only once or twice.

This family of relationships is associated with Zipf’s law.

Voynichese looks surprisingly comfortable in this world.

Its word-like units are not distributed as though every possible token were equally likely.

Common forms recur heavily.

Rare forms form a very long tail.

Older statistical work and later corpus studies have repeatedly noted that Voynich word frequencies look broadly Zipf-like and that its apparent vocabulary can be as varied as language corpora of similar length.

That sounds like a major result.

It is.

It is also extremely easy to overread.

Zipf-like distributions occur beyond natural language.

Generative processes can create them.

Neural language models can reproduce Zipf and Heaps-like laws.

Other communication systems and complex processes can show heavy-tailed rank-frequency behaviour.

And in Voynich, we do not yet know whether a space-delimited token is actually a lexical word.

So the most useful question is not:

Does Voynichese obey Zipf’s law?

It is:

Which frequency and vocabulary laws survive when we change the unit, the manuscript region, the transcription, the comparison corpus and the mechanism?

That turns a familiar graph into a real constraint.


Quick Read

One-sentence answer: Voynich space-delimited tokens show several word-level statistical behaviours associated with natural language—including heavy-tailed rank-frequency distributions, substantial vocabulary growth and roughly ordinary word-level information diversity—but those properties are not uniquely linguistic, depend on how tokens are defined, and coexist with unusually constrained character-level structure, so they support organised information without proving ordinary plaintext language.

  • Zipf’s law describes the tendency for token frequency to decline roughly with frequency rank.
  • Natural-language corpora commonly show Zipf-like rank-frequency curves.
  • Voynich word-like tokens have long been reported to show broadly similar heavy-tailed frequency structure.
  • A Zipf-like curve is evidence of non-uniform organisation, not a unique fingerprint of language.
  • The same law can emerge in other complex systems and in generated text.
  • Vocabulary growth asks how the number of distinct token types rises as more tokens are observed.
  • Heaps’ law is the familiar sublinear vocabulary-growth relationship seen in many natural-language corpora.
  • For Voynich, vocabulary-growth curves and word entropy can look strikingly language-like under common tokenisations.
  • René Zandbergen’s matched-length comparisons found Voynich word-type diversity broadly comparable with Italian and Latin texts even though the visible tokens are shorter and character beginnings are more constrained.
  • That creates a real tension: low character-level freedom and high word-level diversity coexist.
  • A large fraction of Voynich word types occur only once—hapax legomena—especially in specialised label registers.
  • High hapax rates can indicate an open vocabulary, morphology, names, identifiers, transcription noise or productive generation; they do not identify the cause by themselves.
  • Labels have a much flatter frequency distribution than running prose, so manuscript text role materially changes frequency law behaviour.
  • Currier A and B have different frequent tokens and vocabulary preferences, so pooling them can create a synthetic whole-manuscript distribution.
  • Word boundaries and glyph segmentation change the type inventory, frequency ranks and vocabulary-growth curve.
  • A strong theory should reproduce the entire multi-scale profile—characters, tokens, vocabulary growth, local distribution and long-range organisation—not merely one Zipf plot.

The headline is therefore subtle.

Voynich “words” behave enough like real words to matter.

They do not behave so uniquely like real words that the manuscript’s nature is settled.


What Zipf’s Law Actually Says

Take a long English text.

The most frequent words appear extraordinarily often.

Words ranked lower appear progressively less often.

In its classic simplified form, Zipf’s law says that frequency is approximately inversely related to rank.

The second-ranked word occurs about half as often as the first.

The tenth roughly one tenth as often.

Real corpora do not obey that cartoon perfectly.

The exponent varies.

Different ranges of the curve can behave differently.

Corpus size, genre, morphology and unit definition all matter.

Zipf is best understood as a family of heavy-tailed scaling behaviour rather than a magical straight line every language must draw.

This caution becomes even more important for an unknown script.

We are not merely estimating an exponent.

We are also choosing what counts as a word.


The Voynich Rank-Frequency Curve Is Interesting Because the Vocabulary Is Not Flat

Voynichese has highly frequent forms.

Forms such as daiin, ol, chedy and other familiar tokens recur many times under common EVA-based transliterations.

Then frequency drops through a large middle class.

Finally comes a long tail of token types seen once.

This is not what we would expect from a simple generator that chooses independently and uniformly from a fixed dictionary.

Something creates preference.

Some forms are structurally privileged.

Some are rare.

The key scientific question is what mechanism creates the hierarchy.

  • Natural lexical frequency.
  • Morphology.
  • Technical vocabulary.
  • Cipher-state probabilities.
  • Abbreviation frequency.
  • Copy-and-modify generation.
  • A mixture of manuscript registers.

The curve tells us there is a hierarchy.

It does not name the hierarchy.


Why “Zipf-Like Therefore Language” Is Too Fast

Zipf-like laws are strongly associated with language.

They are not exclusive to language.

Power-law and heavy-tailed rank distributions arise in many complex systems.

Machine-generated texts can learn or reproduce them.

Some stochastic models generate approximate Zipf curves.

A human creating pseudo-words by repeatedly modifying common forms can produce a few hubs and a long tail.

A cipher can preserve the frequency hierarchy of plaintext while altering visible forms.

A codebook can create highly uneven code use.

Therefore the correct inference is weaker:

Voynich token frequency is compatible with language-like organisation and inconsistent with many naive random models.

That is already useful.

It does not force plaintext language.


Heaps’ Law Asks a Different Question

Zipf asks how frequencies are distributed among types.

Heaps asks how the vocabulary grows as the text grows.

Read the first 100 words of a novel.

You encounter many new word types.

Read the next 100.

You still discover new words, but many repeat words you already know.

As corpus size increases, vocabulary normally continues to grow sublinearly.

This broad relationship is called Heaps’ law.

It is attractive for Voynich because it asks whether the apparent vocabulary remains productively open rather than saturating quickly like a small fixed codebook.

But the same warning applies.

Vocabulary growth depends directly on tokenisation, spelling variants, transcription resolution and productive generation.

Before asking whether Voynich follows Heaps’ law, we must ask whether our definition of “new word” is stable.


Voynich Vocabulary Keeps Growing

Voynich does not behave like a tiny closed vocabulary repeated endlessly.

As more text is sampled, new token types continue to appear.

René Zandbergen’s matched-length work compared the growth of word types in several Voynich regions with Dante’s Italian and Pliny’s Latin.

The surprising result is that Voynich words can be as varied as those natural-language controls even though Voynich words are visibly shorter and their first characters are much more constrained.

That is a genuine structural puzzle.

If the alphabet strongly limits word beginnings, how does the system keep producing a large vocabulary?

The answer appears to involve greater diversity later in the token.

In matched analyses, Voynich begins with less information than Latin or Italian but catches up through later positions so that whole-token information diversity is surprisingly ordinary.

Voynich compresses choice at the front of a token and restores diversity later.

That pattern is more diagnostic than a Zipf curve by itself.


Word Entropy Creates the Same Paradox

Word entropy asks how uncertain we are about the identity of a whole token before we see it.

A corpus dominated by ten repeated words has low word entropy.

A corpus with a broad, uneven but rich vocabulary has more.

Voynich word entropy has been estimated around the range expected for normal language in long coherent samples.

This sits beside the low character-level conditional entropy discussed elsewhere in the series.

So the manuscript combines:

  • very constrained character transitions;
  • short tokens;
  • substantial whole-token diversity.

A correct mechanism must produce all three.

A model that explains only low entropy is incomplete.

A model that explains only normal word diversity is incomplete.

The tension between them is one of the manuscript’s strongest multi-scale constraints.


Hapax Legomena Are the Long Tail Made Visible

A hapax legomenon is a type that occurs exactly once in the corpus being measured.

Natural language produces many hapax words.

Names.

Rare technical terms.

Productive inflections.

Unusual spellings.

Voynich also has a large hapax tail.

But the percentage depends on which subset and which transcription is counted.

The label article provides a particularly clean example.

Among 299 zodiac labels in one modern analysis, 238 are hapax—about 80% of the label tokens.

That is an extraordinarily flat specialised register compared with ordinary Voynich prose.

It tells us that “hapax rate” is not one manuscript constant.

Document role changes it radically.


High Hapax Does Not Automatically Mean Names

Labels with many unique forms look name-like.

A diagram may indeed need one identifier per figure.

But high uniqueness can be produced by other systems.

  • coordinates;
  • numbers;
  • serial codes;
  • productive morphology;
  • copy-and-modify generation;
  • fine transcription distinctions.

Therefore the correct conclusion is register-level.

Voynich labels use a flatter and more unique token population than running prose.

“They are names” remains one hypothesis among several.


Type–Token Ratio Is Useful—and Dangerous

The type–token ratio divides the number of distinct types by the number of observed tokens.

A text in which every token is unique has a ratio of one.

A text repeating a tiny vocabulary has a much lower ratio.

The problem is sample length.

As a corpus becomes longer, repeated common words accumulate and raw type–token ratio typically falls.

Comparing a 300-token Voynich label set with an 8,000-token Latin passage by raw TTR is therefore misleading.

Matched lengths, moving-window ratios or explicit vocabulary-growth curves are safer.

This is one of the recurring lessons of Voynich statistics:

a number becomes evidence only after the denominator, sample length and comparison process are controlled.


Currier A and B Have Different Frequency Landscapes

Whole-manuscript Zipf plots can hide a major fact.

Voynich is not textually uniform.

In Currier A, daiin-type forms are particularly prominent.

In Currier B, chedy-type families rise dramatically, and some common B words scarcely appear in A.

Pooling A and B therefore creates a mixture distribution.

Mixtures can produce heavy tails even when each component has a different internal law.

This does not invalidate whole-manuscript frequency analysis.

It changes the question.

We should estimate:

  • within-A frequency laws;
  • within-B frequency laws;
  • within-document-role laws;
  • then the mixture.

If the same scaling relation survives each regime, it is a deeper property.

If it appears only after pooling, it may be a mixture artefact.


Labels and Prose Should Not Share One Frequency Law by Default

A label corpus is performing a different job.

Running prose can reuse function words constantly.

A label set may contain one unique identifier per visual element.

If labels are names, coordinates or codes, their frequency law should differ naturally.

This is exactly what the Voynich label data show.

Zodiac and pharmaceutical label populations are much flatter and more hapax-heavy than ordinary running text.

That means document role is not a nuisance to average away.

It is one of the hidden variables controlling the frequency system.

A whole-book vocabulary model should therefore include register explicitly.


Word Boundaries Can Manufacture New Types

The Spaces and Word Boundaries article established that visible gaps are real but their linguistic level is unresolved.

That uncertainty flows directly into vocabulary statistics.

Move one uncertain space.

Two common tokens become one rare token.

Or one rare token becomes two common ones.

The effects propagate.

  • type count changes;
  • hapax count changes;
  • rank changes;
  • Zipf slope changes;
  • vocabulary growth changes;
  • word entropy changes.

The strongest frequency laws should therefore survive reasonable tokenisation alternatives.

If a claimed language law depends on one disputed space convention, it is not yet a robust manuscript property.


Glyph Segmentation Matters Too

Word frequency seems as though it should be independent of alphabet size.

Not entirely.

If two visually different rare glyph forms are collapsed as allographs, two token types may merge.

If one compound is split differently, apparent token identity can change.

Fine-grained v101-like representations can distinguish more word forms than coarser regularised alphabets.

This means the long tail partly reflects palaeographic resolution.

A genuine lexical tail should not disappear when reasonable allographs are normalised.

Frequency-law analysis therefore belongs downstream of the alphabet problem, not outside it.


Local Generation Can Produce a Long Tail

A copy-and-modify process creates new types naturally.

Start with a common token.

Change one glyph.

The result may occur once.

Use that result as another seed.

Now a branching family emerges.

Frequently selected ancestors remain common.

Recent variants remain hapax.

This mechanism can create:

  • heavy-tailed frequencies;
  • ongoing vocabulary growth;
  • many near-neighbour types;
  • many singletons.

Therefore a language-like type distribution does not eliminate local generation.

The generator must then explain higher-level structure—Currier regimes, labels, long-range organisation and page roles.

Frequency law alone cannot decide.


A Cipher Can Preserve or Transform Frequency Laws

A simple substitution preserves plaintext word frequencies exactly if word boundaries and spellings are preserved.

A homophonic or verbose cipher can spread one plaintext unit across several visible forms.

That flattens frequencies and increases visible vocabulary.

A stateful cipher can create families of related ciphertext tokens from one underlying word.

The 2025 Naibbe work is important here because it constructively demonstrates that a historically plausible hand-executable cipher can encrypt meaningful Latin and Italian while reproducing many Voynich-like statistical properties at once.

Naibbe does not solve the Voynich Manuscript.

It changes what counts as a fair null hypothesis.

“Ciphertext would not look like this” is no longer enough.

A cipher model should be compared constructively on the same frequency, entropy and vocabulary-growth metrics.


Natural Language Must Explain the Strange Multi-Scale Combination

Language remains a serious model.

Word-level diversity supports it.

Long-range organisation supports it.

Distributional classes and topic-like locality support it.

But natural language also inherits hard questions.

  • Why is conditional character entropy unusually low?
  • Why are token beginnings so restricted?
  • Why are exact word bigrams comparatively weak?
  • Why do line boundaries affect token form so strongly?
  • Why do near-neighbour families dominate the visible vocabulary?

A natural-language solution may invoke abbreviation, unusual orthography or encoding.

Those are legitimate possibilities.

They must reproduce the whole statistical stack.

Zipf is one floor tile, not the building.


The 2013 PLOS Studies Show Why Multiple Statistics Matter

Two influential 2013 studies approached Voynich from different statistical directions.

Montemurro and Zanette found complex long-range word organisation compatible with real-language sequences and extracted location-informative token networks.

Amancio and colleagues compared Voynich with natural-language corpora and shuffled controls using first-order word statistics, complex networks and intermittency measures.

Their overall result was that Voynich differed strongly from simple shuffled word sequences and was mostly compatible with natural-language behaviour under the tested measures.

That is substantial evidence against naive randomness.

It is not proof of plaintext.

A structured cipher can preserve language-derived organisation.

A sophisticated generator can be non-random by design.

The lesson from those papers is methodological:

when one statistic is non-diagnostic, cross several independent statistics and compare mechanisms, not labels.


A 2026 Preprint Pushes the Unit Problem Even Harder

A very recent 2026 preprint by Liudmila Rozanova and Alexander Temerev argues that the usual assumptions “glyph = letter”, “token = word” and “blank = word space” should all be earned rather than granted.

The authors compare Voynich with prose, cipher and pseudo-text controls and report an open, hapax-rich visible vocabulary alongside weak exact token-order dependence and stronger edge-level structure.

This is new work and should be treated as a preprint, not settled consensus.

Its broad methodological challenge is nevertheless useful:

frequency laws are only as linguistically interpretable as the units to which we apply them.

That principle is independent of whether the paper’s specific conclusions survive future scrutiny.


The Best Comparison Is Not “Voynich Versus English”

English is easy to obtain.

It is not a neutral default.

Languages differ in morphology.

A highly inflected language produces more surface word types than an analytic language.

Genres differ in vocabulary.

A herbal or technical register differs from narrative prose.

Historical spelling multiplies variants.

Abbreviation changes visible word length.

Therefore frequency-law claims should be tested against matched families of controls.

The fourth article in this batch owns that control problem fully.


What Survives the Frequency-Law Work

  • Voynich token frequencies are strongly non-uniform and heavy-tailed.
  • The rank-frequency distribution is broadly Zipf-like under common word-token definitions.
  • Voynich apparent vocabulary continues to grow substantially with corpus size.
  • Whole-token diversity and word entropy can be comparable with natural-language texts of similar length.
  • This coexists with unusually constrained character-level predictability and short visible tokens.
  • Many token types are rare or unique.
  • Labels are especially hapax-heavy and have a much flatter frequency distribution than running prose.
  • Currier A/B, document role, transcription and tokenisation materially affect the frequency landscape.
  • Zipf-like and vocabulary-growth behaviour are compatible with natural language but not exclusive to it.
  • Ciphers and generated systems can reproduce important parts of the same profile.
  • The strongest evidence comes from the combination of frequency laws with independent structural and long-range constraints.

What Does Not Survive as Established Knowledge

  • Zipf’s law proves natural language.
  • A Zipf-like slope identifies the Voynich language.
  • Heaps-like vocabulary growth proves a lexicon rather than productive generation.
  • A high hapax rate proves proper names.
  • A high type–token ratio proves semantic richness.
  • Whole-manuscript frequency laws can be interpreted without separating Currier regimes.
  • Labels should follow the same vocabulary distribution as prose.
  • Visible spaces are proven lexical boundaries.
  • Every rare transcription form is a distinct lexical type.
  • Language-like word entropy resolves the unusual character entropy problem.

The statistical resemblance survives.

The exclusive diagnosis does not.


A Better Frequency-Law Analysis

  1. Define the tokenisation before counting.
  2. Report the transcription and glyph regularisation level.
  3. Measure rank-frequency curves, not only one fitted exponent.
  4. Measure vocabulary growth across matched token lengths.
  5. Report hapax counts and type–token measures with sample-size controls.
  6. Separate Currier A, Currier B and mixed pages.
  7. Separate labels, prose, circular text and record-like entries.
  8. Test sensitivity to uncertain spaces.
  9. Compare against natural languages with different morphological systems.
  10. Compare against historically plausible ciphers and explicit generators.
  11. Require one model to reproduce both word-level diversity and character-level constraints.
  12. Validate the fitted laws on held-out bifolia or manuscript regions.

What Would Count as a Real Frequency-Law Breakthrough?

Imagine three competing mechanisms are implemented fully.

  • A natural language with historically plausible abbreviation.
  • A hand-executable fifteenth-century cipher.
  • A constrained local generator.

Each produces corpora matched to Voynich in length, section size and document roles.

Researchers compare:

  • Zipf curves;
  • vocabulary growth;
  • hapax rates;
  • word entropy;
  • character entropy;
  • near-neighbour family density;
  • token-order dependence;
  • long-range locality;
  • label/prose register splits.

One mechanism consistently reproduces the whole profile across unseen pages with fewer free parameters.

That would be much more informative than fitting one power law.

The breakthrough would come from the joint distribution of many constraints.

The question is not whether Voynich looks like language in one graph. It is whether one plausible mechanism keeps looking like Voynich no matter which graph we draw.


Primary School: The Few Common Words and the Many Rare Ones

Ask a child to count the words in one page of a story.

Words such as “the” and “and” appear many times.

Some names appear once.

Make a bar chart.

The child discovers that language does not use every word equally.

Then explain that Voynich has a similar unequal hierarchy.

The important next question is whether only language can create it.


Lower Secondary: Grow the Vocabulary

Read a text in blocks of 100 words.

After each block, count how many different words have appeared so far.

Plot corpus size against vocabulary size.

The line keeps rising but slows.

Now compare a fixed codebook generator.

Its vocabulary saturates.

Then compare a mutation generator.

Its vocabulary may continue growing too.

The exercise shows why vocabulary growth is informative but not uniquely linguistic.


Upper Secondary: Change the Tokenisation

Give students the same artificial text with ambiguous spaces.

Version A treats each ambiguous gap as a word boundary.

Version B merges across those gaps.

Calculate:

  • number of types;
  • number of hapax;
  • frequency ranks;
  • type–token ratio.

Every statistic changes even though the ink pattern is unchanged.

This is the Voynich unit problem in quantitative form.


JC and Adult Readers: Scaling Laws Are Model Diagnostics

At a higher level, scaling laws should be treated as signatures generated by mechanisms.

A proposed mechanism implies a distribution over ranks, vocabulary growth and singleton rates.

Instead of asking whether an observed curve resembles a textbook law, estimate how probable the Voynich curve is under competing generative models.

The comparison should condition on:

  • corpus length;
  • section mixture;
  • tokenisation;
  • morphological richness;
  • historical spelling;
  • abbreviation;
  • cipher transformations.

Scaling then becomes one dimension of model selection rather than a yes/no language detector.


A Parent and Teacher Guide

  1. Teach token and type before Zipf.
  2. Show that common words and rare words coexist naturally.
  3. Measure vocabulary growth rather than only total vocabulary.
  4. Explain hapax as “seen once”, not “must be a name”.
  5. Keep sample lengths equal when comparing corpora.
  6. Change tokenisation and watch the statistics move.
  7. Compare language with cipher and generation controls.
  8. Never let one attractive graph decide an entire manuscript.

The wider lesson is fundamental:

a statistical law is strongest when it distinguishes mechanisms, not when it merely resembles a familiar example.


Reader Checklist: Before You Say “Voynich Follows Zipf, Therefore…”

  1. What counts as a token?
  2. Which spaces are treated as boundaries?
  3. Which transliteration alphabet is used?
  4. Are rare glyph variants regularised?
  5. What corpus length is being analysed?
  6. Are Currier A and B pooled?
  7. Are labels mixed with prose?
  8. Is the full rank-frequency curve shown or only one fitted exponent?
  9. How does vocabulary size grow with text length?
  10. How many hapax occur at matched sample sizes?
  11. What natural-language controls are used?
  12. What cipher controls are used?
  13. What generation controls are used?
  14. Does the model also explain low character entropy?
  15. Does it explain exact repetition and weak token bigrams?
  16. Does the result survive unseen bifolia or sections?

Frequently Asked Questions

Does Voynichese follow Zipf’s law?

Its space-delimited token frequencies show broadly Zipf-like heavy-tailed behaviour under common transliterations. The exact fit and parameters depend on corpus and unit choices.

Does that prove Voynichese is a natural language?

No. Zipf-like distributions are common in language but are not exclusive to it. Structured generation and ciphers can reproduce similar frequency hierarchies.

What is Heaps’ law?

It describes the sublinear growth of vocabulary size as more tokens are observed. It is common in natural-language corpora but can also emerge in generated systems.

Does Voynich have a large vocabulary?

Under common tokenisations, it has substantial and continuing word-type diversity. Matched-length studies show whole-token diversity comparable with ordinary language corpora despite shorter and more constrained visible words.

What is a hapax legomenon?

A token type that occurs exactly once in the measured corpus.

Does Voynich have many hapax?

Yes, although the proportion depends strongly on corpus definition. Specialised labels are particularly singleton-heavy; one modern zodiac-label analysis finds about 80% of the 299 labels occur only once.

Why is Voynich word diversity surprising?

Because visible character sequences are unusually constrained and tokens are short, yet the corpus still produces a broad variety of whole-token forms. Any complete mechanism must explain both facts simultaneously.

Could a cipher preserve Zipf-like behaviour?

Yes. Depending on the cipher, plaintext frequency hierarchy can be preserved or transformed. Modern constructive work such as the Naibbe cipher shows that meaningful plaintext can be converted into ciphertext with multiple Voynich-like statistical properties.

What is the strongest current conclusion?

Voynich apparent words have a rich, non-random frequency and vocabulary structure compatible with language-like organisation. The structure is real; the linguistic identity of the units remains unproven.


Related eduKateSG Reading


Research and Further Reading


The Final Idea

Voynichese passes an important test and fails an easy conclusion.

Its vocabulary is not a flat pile of arbitrary strings.

Some forms dominate.

Others form a deep middle.

A long tail appears once.

New forms keep arriving as the corpus grows.

At whole-token scale, this can look remarkably ordinary.

Then we look inside the tokens.

The character system becomes unusually restrictive.

We look between tokens.

Exact word order is weaker than expected.

We look across pages.

Vocabulary becomes local, Currier-sensitive and register-sensitive.

This is why the manuscript remains difficult.

Every scale gives us real structure.

No single scale tells us what the structure is for.

The right response to a language-like graph is not “solved”. It is to ask which mechanism can produce all the other graphs at the same time.


Continue Through the Voynich Research Map

This article is one specialist node in eduKateSG’s larger Voynich research library. Return to the canonical master to see the physical, visual, textual and historical evidence in one continuous argument.

Frequency resemblance is useful only when the comparison unit, corpus length, transcription and competing mechanisms are controlled.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading