VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Word-Length Problem: Why Voynich Words Keep Coming Out About the Same Size

Voynich words have a strange habit.

They keep coming out about the same size.

Not exactly.

There are short tokens.

There are long ones.

But the apparent vocabulary does not spread across word lengths in the broad, lopsided way many natural-language lexicons do.

Under one influential set of assumptions about what counts as a character and what counts as a word, Jorge Stolfi found that the distribution of distinct Voynich word types by length is almost perfectly symmetric around a mean of roughly 5.5 symbols.

The shape is strikingly close to a binomial distribution.

That is not a small decorative curiosity.

A word-length distribution is a compressed fingerprint of how a writing system builds its tokens.

If words can freely accumulate prefixes, stems, suffixes and compounds, length varies one way.

If tokens are assembled from a fixed number of optional slots, length varies another way.

If vowels are omitted, words shorten.

If abbreviation compresses common sequences, words shorten and may cluster.

If a verbose cipher expands source units, words may lengthen.

If visible spaces do not mark lexical words, the entire distribution may be describing the wrong unit.

This makes the word-length problem unusually powerful.

It is also unusually fragile.

Change the alphabet.

A bench becomes one symbol instead of two.

A minim string becomes one compound instead of several strokes.

Move an uncertain space.

Two short words become one long word.

The graph changes without a single mark on the parchment changing.

The Voynich word-length anomaly is real enough to demand explanation, but it is only as real as the character and word definitions used to measure it.


Quick Read

One-sentence answer: under Stolfi’s Currier-style symbol counting and conventional space-delimited tokenisation, the lengths of distinct Voynich word types form an unusually narrow, almost symmetric distribution centred near 5.5 symbols and closely approximated by a binomial curve; this is a strong structural constraint but not a unique signature of cipher, language or generation, and it changes with character segmentation, word boundaries, Currier regime and text role.

  • Word length can be measured over token occurrences or over distinct word types.
  • The two distributions are not the same because frequent words tend to be shorter.
  • Stolfi’s famous result concerns the distribution of distinct word types under a particular symbol definition.
  • He counted EVA ch/sh and pedestal-gallows-like units synthetically, following Currier-style symbol concepts.
  • Under those assumptions, the word-type length distribution is nearly symmetric around about 5.5 symbols.
  • Stolfi noted that the type distributions of main text and labels are surprisingly similar even though their token-frequency distributions differ strongly.
  • He proposed a combinatorial/binomial-style explanation in which structured token forms resemble codes generated from optional positions.
  • That proposal is a model, not an accepted decipherment.
  • Zandbergen calls word-length statistics among the least reliable Voynich statistics because both character and word definitions are uncertain.
  • Reddy & Knight noted that the Voynich length profile resembles devowelled English, Arabic and Chinese represented in Pinyin under some comparisons.
  • This keeps omitted-vowel or consonant-heavy representations viable.
  • Abbreviation can narrow visible word lengths by compressing predictable material.
  • A templatic generator can produce a near-binomial distribution naturally if several positions are independently optional.
  • A cipher can also produce constrained lengths, especially if code groups or verbose substitutions have regular structure.
  • Labels and prose having similar type-length structure may indicate a shared word-construction system despite different vocabularies.
  • A real mechanism should reproduce the shape across unseen sections without being tuned directly to the histogram.

Word length therefore sits at the intersection of the alphabet, vowels, spaces, morphology, abbreviation and generation problems.

It is not one more statistic.

It is a compression of several of them.

First Separate Token Length From Word-Type Length

This distinction is essential.

A token is one occurrence in the manuscript.

If daiin appears one hundred times, that contributes one hundred tokens.

A type is the abstract distinct form.

Those one hundred occurrences contribute one word type.

Natural languages often use short words frequently.

Function words are a classic example.

Therefore the token-length distribution tends to be weighted toward short lengths more strongly than the type-length distribution.

Voynich behaves similarly in this broad respect.

That means a claim such as “Voynich words average 5.5 characters” can be misleading unless it states:

  • tokens or types;
  • which transcription;
  • which symbol definition;
  • which corpus region.

The famous symmetry belongs particularly to the type distribution under Stolfi’s assumptions.

Why 5.5 Symbols Is Such a Strange Centre

A mean near 5.5 is not mysterious by itself.

Many languages have words in that broad range.

The unusual feature is the symmetry and narrowness.

There are relatively few extremely short or extremely long distinct forms compared with what many unconstrained lexicons produce.

The graph rises toward a central region and falls away in a balanced fashion.

That resembles the outcome of repeated binary choices.

For example, imagine ten optional positions.

Each position independently appears or disappears with some probability.

The number of occupied positions follows a binomial-like distribution.

Most outputs have middling length.

Very short and very long outputs are rare.

This is why Stolfi explored a combinatorial code-style explanation.

The graph looks as though Voynich words may be generated from constrained positions rather than assembled from unconstrained character strings.

A Binomial-Looking Histogram Is Not Proof of a Binary Code

This is the first firewall.

Many mechanisms can produce a bell-like or near-binomial length distribution.

Independent optional slots are one.

Strong morphology with bounded affix count can be another.

Abbreviation can compress tails.

A transcription system can regularise complex units.

A cipher can enforce group-length rules.

A copy-and-modify generator can stay close to common seed lengths.

Therefore the shape is a constraint.

It does not name the mechanism.

The winning explanation must also reproduce character ordering, word families, Currier differences, labels, entropy and syntax.

Stolfi’s Symbol Definition Matters

Stolfi did not count every EVA code point as one independent symbol.

He followed a more synthetic Currier-like concept for several units.

EVA ch and sh count as one symbol.

Complex pedestal-gallows forms are treated as units.

This choice shortens many visible tokens relative to a purely analytical EVA character count.

Use another alphabet and the mean changes.

The shape can change too.

This is why Zandbergen warns that word-length statistics are particularly vulnerable to representation.

The manuscript has not changed.

Our measurement unit has.

“Five symbols long” is not a purely visual fact until we have decided which strokes belong to one functional symbol.

The Alphabet Problem Sits Under the Length Problem

The Alphabet Problem asked how many functional characters Voynich actually has.

The length problem asks how many of those characters fit into one word-like unit.

If a bench is one character, one token has length 5.

If the bench is two components, the same token has length 6.

If a pedestalled gallows is decomposed further, perhaps length 7.

Multiply those differences across thousands of word types and the entire histogram moves.

A robust mechanism should therefore explain why the unusual narrowness survives—or fails to survive—reasonable alphabet models.

If only one arbitrary segmentation creates the binomial shape, the anomaly belongs partly to the transcription.

If many plausible segmentations preserve it, the manuscript-level constraint strengthens.

The Space Problem Sits Under It Too

Word length assumes we know where words begin and end.

Voynich spaces are visible but not proven lexical boundaries.

Some gaps are uncertain.

Merge across one gap and two short tokens become one long token.

Split one dense group and a long token becomes two short ones.

These operations change exactly the tails of the length distribution that make Stolfi’s symmetry so striking.

This creates an important sensitivity test.

Recompute the histogram under several plausible gap thresholds.

Does the narrow shape survive?

If yes, it is robust.

If no, “word length” may be partly a spacing convention rather than lexical design.

Labels Create a Remarkable Comparison

Stolfi observed something especially interesting.

The distribution of distinct word lengths in label text resembles that of the main text surprisingly well.

This happens even though labels and running prose behave very differently in frequency.

Labels are flatter.

More singleton-heavy.

They avoid q-like openings.

Yet their available type lengths can inhabit a similar structural envelope.

This suggests a shared word-construction mechanism across document roles.

The lexicon changes.

The register changes.

The underlying constraints on how long a legal form can be may remain similar.

That is a powerful clue.

Shared Length Structure Does Not Prove Shared Semantics

A naming system and prose language can use the same orthography.

A code generator can use the same token template for labels and paragraphs.

A cipher can encode both under one group grammar.

Therefore similar type-length distributions tell us about construction more directly than meaning.

This is exactly the kind of cross-register invariant a good model should explain.

If the same mechanism produces legal token envelopes in prose and labels while allowing their frequency profiles to differ, it gains strength.

Reddy & Knight’s Devowelled Comparisons Matter

Reddy and Knight approached the length question from another direction.

They noted that Voynich word-length behaviour resembles certain representations in which vowels are absent or represented differently.

Devowelled English is one useful control.

Arabic provides another consonant-centred comparison.

Chinese written in Pinyin without ordinary assumptions offers another structural comparison.

The point is not that Voynich is English without vowels, Arabic or Chinese.

The point is that removing or reorganising vowel information changes word-length distributions substantially.

This links the Vowel Problem directly to the Length Problem.

If Voynich suppresses predictable vowels, its narrow visible lengths become less anomalous.

Omitted Vowels Make a Strong Prediction

An omitted-vowel hypothesis should do more than shorten words.

It should recover an underlying language.

After vowel restoration:

  • word lengths should become historically plausible;
  • consonant skeletons should map consistently to lexical families;
  • grammar should constrain which vowels are possible;
  • the restored distribution should resemble the proposed source language under comparable spelling.

If every skeleton can be vocalised dozens of ways and the translator chooses whichever yields a useful word, the length explanation has purchased flexibility rather than evidence.

A real omitted-vowel mechanism reduces ambiguity through morphology and syntax.

Abbreviation Can Narrow Word Length From the Other Direction

Medieval abbreviation compresses longer underlying forms into shorter visible ones.

Frequent suffixes become one sign.

Internal material disappears under contraction.

Predictable endings are suspended.

Long source words can converge toward a narrower visible length range.

This makes abbreviation one of the most plausible historical mechanisms capable of changing the histogram.

But compression should be systematic.

Expand the abbreviations under fixed rules.

Does the source-side length distribution become more natural?

Do common Voynich endings map to common historical endings?

Does the same expansion work in labels and prose?

If not, abbreviation is only a possibility label.

A Fixed-Slot Template Produces the Shape Naturally

Voynich words often look as though they contain positional zones.

Optional opening material.

A core.

Bench or gallows structures.

An ending.

If several positions are independently optional, length naturally clusters near the number of usually occupied positions.

Extreme lengths require many slots to be simultaneously absent or present and therefore become rare.

This is one route to a binomial-looking distribution.

It fits well with Stolfi’s old “core–mantle–crust” style descriptions of token architecture.

But positional templates occur in both language and generated systems.

Morphology itself can be templatic.

A Semitic-style root-and-pattern system is structured.

A code is structured.

The histogram reveals bounded composition.

It does not tell us whether the composition is linguistic.

Copy-and-Modify Generation Also Tends to Preserve Length

Start with one token of length five.

Copy it.

Replace one character.

Length remains five.

Add or remove one character occasionally.

Most descendants remain close to the original length.

Repeated local mutation therefore generates dense word families with narrow length spread naturally.

This is one reason copy-modify generation can reproduce aspects of Voynich word architecture.

The challenge is global.

Can the same mechanism reproduce Currier states, label registers, long-range topic-like organisation, line effects and the correction profile?

Length preservation alone is not enough.

A Cipher Can Impose Group-Length Constraints

Classical cryptography often works with fixed or semi-fixed groups.

A nomenclator can replace common words with compact codes.

A verbose cipher can encode one plaintext letter as several ciphertext signs.

A homophonic system can choose among groups of similar length.

Word boundaries may even be preserved visually while internal length is transformed.

Therefore an unusual length distribution does not rule out cipher.

A constructive cipher control should encrypt plausible source texts and measure the resulting distribution.

The 2025 Naibbe result is relevant again because realistic hand-cipher behaviour can reproduce several Voynich-like statistics simultaneously.

The key question is whether one cipher architecture reproduces the length histogram as a consequence rather than a parameter tuned directly to it.

Natural Morphology Can Be Narrow Too

It would be a mistake to treat narrow word length as inherently artificial.

Languages differ dramatically in morphology and orthography.

A language with short roots and bounded suffixes can produce a fairly concentrated visible length range.

A consonant-heavy script can compress variation.

A syllabic system can change the number of visible signs per spoken word.

This is why the Control Problem article rejected English as a universal baseline.

The right comparison suite includes languages with different morphological and writing-system structures.

The anomaly is strongest when Voynich remains unusually symmetric after those controls are matched fairly.

Currier A and B Should Have Separate Length Profiles

Whole-manuscript averages can hide regime differences.

Currier A and B have different common words.

B contains many -edy-like families that are rare in A.

A favours other token forms.

These vocabulary differences can change token-length frequencies.

A mechanism-level word-construction constraint should survive both, perhaps with shifted parameters.

If only one regime produces the near-binomial type histogram, the anomaly is not universal.

If A, B and intermediate/RZ-C material all retain the same envelope, the shared construction grammar becomes more significant.

The Drift Problem Adds Another Test

If Voynich vocabulary drifts gradually through production, what happens to word length?

A continuous mutation generator might preserve mean length remarkably well while changing character identities.

A changing language or genre might alter average lengths.

An evolving abbreviation system might gradually compress or expand them.

Therefore length can serve as a second dimension of the drift model.

Does mean length remain stable while vocabulary changes?

Does variance remain stable?

Do the tails change at Currier transitions?

A stable length envelope across changing vocabularies would suggest a persistent higher-level construction rule.

Line Geometry Can Bias the Observed Token-Length Distribution

Page planning creates another confounder.

Text often fits between drawings.

Narrow spaces may favour short tokens or abbreviated variants.

If illustrated pages contain systematically more spatial constraints than text-only pages, token lengths can shift.

A strong analysis should therefore compare:

  • open text regions;
  • image-constrained lines;
  • labels;
  • text-only pages.

If length changes strongly with available width, page geometry is part of the generating process.

If the same narrow type distribution survives unconstrained text, the writing system itself carries more of the effect.

The Almost-Correctionless Script Adds Another Constraint

A complex word-building system must be executable by hand.

Voynich writers produce thousands of tightly constrained forms with few conspicuous corrections.

If the length distribution comes from a ten-slot combinatorial code, could a human maintain it fluently?

If it comes from abbreviation, are the shortening rules easy enough to apply?

If it comes from morphology, does ordinary language competence explain the fluency?

Word-length theories should therefore be judged partly by human production cost.

A beautiful mathematical generator that would require constant counting and checking may conflict with the manuscript’s clean execution.

Why the Similarity Between Text and Labels Is a Particularly Strong Constraint

Suppose running text represents grammatical words.

Suppose labels represent proper names or identifiers.

The two registers have different semantic and frequency demands.

Yet if their distinct type-length distributions remain similar, one higher-level constraint may govern both.

  • same legal slot grammar;
  • same abbreviation envelope;
  • same cipher group construction;
  • same syllabic orthography.

This cross-register invariance is harder for a theory to explain accidentally than one histogram in one section.

It deserves renewed measurement under modern transcription alternatives.

Word-Length Statistics Are a Perfect Example of the Control Problem

Compare Voynich only with English and the distribution may look extraordinary.

Compare with devowelled English and the gap changes.

Compare with Arabic and another part changes.

Compare with an abbreviating medieval text and another.

Compare with a fixed-slot generator and the shape may be easy to reproduce.

The statistic therefore has no meaning without control families.

The best comparison suite should vary one mechanism at a time.

  • full spelling vs devowelled spelling;
  • unabbreviated vs abbreviated manuscript text;
  • plaintext vs simple cipher;
  • plaintext vs verbose cipher;
  • language vs templatic generator;
  • language vs copy-modify generator.

Then the length curve becomes diagnostically useful.

A Good Mechanism Should Predict More Than the Mean

Matching 5.5 symbols is easy.

Set a generator’s average length to 5.5.

Done.

A serious model should predict the entire shape.

  • probability of length 1;
  • length 2;
  • length 3;
  • through the centre;
  • through the long tail.

It should predict token and type distributions separately.

It should predict labels.

Currier regimes.

Line positions.

Unseen bifolia.

One fitted mean is not explanatory compression.

The Distribution of Word Families Should Be Conditioned on Length

Near-neighbour families are central to Voynich.

But edit distance depends partly on length.

If almost every word has five or six symbols, any random pair has more opportunities to align positionally than if lengths vary from one to twenty.

Therefore word-family density should be compared against controls with the same length distribution.

Otherwise narrow lengths can inflate apparent family structure.

This does not make the family relationships disappear.

It improves the baseline.

A strong family anomaly should remain after length is matched.

Entropy Should Be Conditioned on Length Too

Short constrained words naturally have lower character uncertainty than long unconstrained words.

If Voynich length is unusually narrow, part of its entropy profile may reflect that envelope.

Conversely, the internal character grammar may be what creates the narrow lengths.

The two statistics are coupled.

A correct mechanism should generate both simultaneously.

This is another reason individual anomalies should not be counted as independent victories automatically.

Word length, slot grammar, character entropy and word-family density may be several shadows of one underlying construction rule.

What Survives the Word-Length Work

  • Voynich visible word-like units are relatively short and strongly length-constrained.
  • Token-length and type-length distributions must be distinguished.
  • Under Stolfi’s Currier-style symbol definitions, the distinct word-type length distribution is unusually symmetric around roughly 5.5 symbols.
  • The shape is close enough to a binomial distribution to motivate combinatorial/slot-based models.
  • That shape is not uniquely diagnostic of a code or generator.
  • Word-length statistics are highly sensitive to glyph segmentation and space definitions.
  • Labels and main text show surprisingly similar type-length structure under Stolfi’s analysis despite very different token distributions.
  • Reddy & Knight’s comparisons keep omitted-vowel and consonant-heavy representation models plausible.
  • Abbreviation, syllabic writing, templatic morphology, cipher and local generation can all alter length distributions.
  • A valid mechanism should reproduce token/type shapes across independent registers and regimes without directly tuning the histogram.

What Does Not Survive as Established Knowledge

  • Every Voynich word is exactly five or six characters long.
  • The average word length is universally 5.5 under every transcription.
  • The near-binomial distribution proves a binary code.
  • The distribution proves meaningless generation.
  • The distribution proves omitted vowels.
  • The distribution proves Arabic or another specific language.
  • Stolfi’s symbol segmentation is the uniquely correct alphabet.
  • Visible spaces are proven lexical word boundaries.
  • Labels have the same meanings as prose because their type lengths are similar.
  • One matched mean word length validates a decipherment.

The narrow construction envelope survives.

The mechanism that creates it does not.

A Better Word-Length Analysis

  1. Report token and type length distributions separately.
  2. State exactly what counts as one symbol.
  3. Repeat under analytical and synthetic alphabets.
  4. Test several plausible word-boundary treatments.
  5. Separate Currier/RZ regimes.
  6. Separate prose and labels.
  7. Compare image-constrained and text-only regions.
  8. Use matched historical language controls with different morphology.
  9. Include devowelled and abbreviated controls.
  10. Include constructive cipher and structured-generation controls.
  11. Match length distributions before testing edit-distance word families.
  12. Ask one mechanism to reproduce length, entropy and family structure together.
  13. Test the fitted model on unseen bifolia.

What Would Count as a Real Word-Length Breakthrough?

Imagine several independently plausible character segmentations are tested.

The same narrow type-length envelope survives them all.

Uncertain spaces are varied and the effect remains.

A compact six-slot construction grammar is then learned from one subset of the manuscript.

Without fitting length directly, it predicts the word-length distribution of held-out prose, labels and Currier regimes.

The same slots correspond to independently discovered character functions.

A decipherment then maps those slots onto stable morphological, syllabic or cipher operations.

The model also predicts why the correction profile is so clean and why word families cluster by single edits.

That would make the histogram evidence for a real generative grammar rather than a visual curiosity.

The word-length problem will be solved when the narrow histogram stops being a statistic we fit and becomes a necessary consequence of a mechanism that explains the rest of Voynichese too.

Primary School: Optional Pieces Make a Bell Shape

Give a child five optional beads.

For each bead, flip a coin to decide whether it goes on a string.

Repeat many times.

Very short and very long strings are rare.

Middle-sized strings are common.

This gives an intuitive picture of why optional token slots can create a binomial-like length distribution.

Lower Secondary: One Word, Two Alphabets, Two Lengths

Write an invented compound symbol.

In Alphabet A, count the compound as one character.

In Alphabet B, count its two components separately.

The same ink has two measured lengths.

Students see why Voynich word-length statistics depend on the character model.

Upper Secondary: Remove the Vowels

Take a paragraph of ordinary English.

Measure word lengths.

Now remove most vowels and measure again.

The distribution shifts and narrows.

This demonstrates why word-length resemblance to devowelled controls is relevant without proving that Voynich actually omits vowels.

JC and Adult Readers: Length as a Marginal of a Generative Grammar

At a higher level, word length is not an independent phenomenon.

It is the marginal distribution produced by a hidden token grammar.

If a grammar has slots with occupancy probabilities, those probabilities imply a length distribution.

If a morphological language has stem and affix distributions, those imply another.

If a cipher maps source units to groups, its mapping distribution implies another.

The correct model should therefore predict length from deeper rules rather than model length independently.

That is explanatory direction.

A Parent and Teacher Guide

  1. Separate token length from distinct type length.
  2. Ask what counts as one Voynich character.
  3. Ask what counts as one Voynich word.
  4. Treat 5.5 symbols as representation-dependent rather than universal.
  5. Use the binomial shape as a clue to construction, not proof of code.
  6. Compare omitted-vowel, abbreviation, language, cipher and generation controls.
  7. Look for the same envelope in labels and prose.
  8. Make the mechanism predict the histogram rather than tuning itself to it.

The transferable lesson is:

a measurement can be mathematically exact and still answer the wrong question if the units being counted have not been established.

Reader Checklist: Before You Use Voynich Word Length as Evidence

  1. Are you measuring tokens or types?
  2. Which transcription is used?
  3. What counts as one character?
  4. Are benches counted synthetically or analytically?
  5. How are minim strings counted?
  6. How are rare glyph variants regularised?
  7. What counts as a space?
  8. Are uncertain spaces included?
  9. Are Currier regimes pooled?
  10. Are labels mixed with prose?
  11. Is page geometry affecting token choice?
  12. What natural-language controls are used?
  13. Are devowelled controls used?
  14. Are historical abbreviation controls used?
  15. Are cipher and generator outputs compared?
  16. Does the mechanism predict the full distribution, not only the mean?
  17. Does it generalise to unseen bifolia?

Frequently Asked Questions

How long is a typical Voynich word?

The answer depends on transcription and whether one measures tokens or distinct types. Under Stolfi’s Currier-style symbol definition, the distinct word-type distribution is centred around roughly 5.5 symbols.

Why is the distribution unusual?

The distinct word-type lengths are unusually narrow and close to symmetric under that representation, approximating a binomial shape more closely than many ordinary natural-language lexicons.

Does that prove the text was generated?

No. Templatic generation can produce the shape naturally, but so can bounded morphology, omitted-vowel writing, abbreviation and some cipher mechanisms.

Why does character segmentation matter?

If a compound glyph is counted as one symbol in one alphabet and several in another, every affected word changes measured length. The histogram is therefore partly representation-dependent.

Why do spaces matter?

Moving one uncertain word boundary can replace two short tokens with one long token, directly altering the length distribution.

Why are devowelled languages relevant?

Removing or representing vowels differently compresses visible word lengths. Reddy & Knight found structural similarities that keep omitted-vowel models worth testing without identifying a specific language.

What did Stolfi conclude?

He documented the unusual near-binomial type-length distribution and explored a combinatorial code-like model capable of producing it. That model remains an explanatory proposal, not a decipherment.

What is the strongest current conclusion?

Voynich word-like forms occupy a strongly constrained length envelope whose exact numerical shape depends on character and boundary definitions. Any serious language, cipher or generation model must reproduce that envelope alongside the manuscript’s other structures.

Related eduKateSG Reading

Research and Further Reading

The Final Idea

The word-length problem is a beautiful trap.

The graph looks mathematical.

Clean.

Almost too clean.

A narrow symmetric hill centred near 5.5 symbols.

It invites a mechanism immediately.

Binary code.

Fixed template.

Generated text.

Then we remember that the x-axis itself depends on decisions we have not fully solved.

What is one character?

What is one word?

What does a space mean?

Once those uncertainties are made visible, the graph does not become useless.

It becomes better science.

Because a real mechanism will have to survive several reasonable ways of measuring the same ink.

The Voynich words keep coming out about the same size. The real question is whether the manuscript built them that way—or whether our current way of cutting the manuscript into words and characters makes them look that way.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading