VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | Cipher, Plaintext or Generated System?

There is a question hiding underneath almost every Voynich argument.

Not:

What does this word mean?

But:

What kind of thing is this writing system doing?

If the visible text is ordinary language in an unfamiliar script, one kind of solution follows.

If it is heavily abbreviated language, another follows.

If it is ciphertext, another.

If it is a constructed notation, another.

If it is generated pseudo-text, another.

And if several mechanisms are layered together, the word “cipher” may be simultaneously too simple and partly right.

This article does not choose a winner.

It compares the mechanism families.

Quick Read

  • Voynichese is strongly structured and non-random in ordinary senses, but structure alone does not identify its mechanism.
  • A natural-language model explains why words, local vocabulary, syntax-like preferences and document organisation exist, but must explain unusually strong glyph and word constraints.
  • An abbreviation model is historically plausible because medieval scribes routinely compressed language, but it must account for Voynich-specific regularities across thousands of tokens.
  • A cipher model remains viable. A 2025 Cryptologia study showed that a historically plausible hand-operable verbose homophonic substitution cipher can encrypt Latin and Italian while reproducing many Voynich-like statistical properties.
  • That Naibbe result demonstrates compatibility with a cipher family; it does not identify the actual Voynich mechanism.
  • Generated-text and hoax models can reproduce some low-level properties, proving that “looks language-like” is not sufficient evidence for meaningful plaintext.
  • Generated models face a harder burden at larger scale: page families, local vocabulary, labels, Currier regimes, scribal production, visual relationships and manuscript-wide organisation.
  • Constructed-language or specialised-notation models remain possible but require their own predictive mechanism rather than functioning as a catch-all.
  • Hybrid models are historically plausible: natural language can be abbreviated, encoded, supplied with nulls, transformed orthographically and mixed with specialist notation.
  • The correct question is not “Which theory can imitate one Voynich feature?” but “Which mechanism explains the largest set of independent features with the fewest exceptions?”

The First Mistake: Treating the Visible Text as the Underlying Language

Suppose the underlying message were Latin.

That does not mean the visible text should look statistically like ordinary Latin.

An encoding layer can radically alter surface behaviour.

Abbreviation can shorten common sequences.

Homophonic substitution can spread one plaintext letter across several ciphertext forms.

Nulls can introduce symbols with no direct plaintext value.

Verbose ciphers can expand one letter into several glyphs.

Specialised orthography can create constraints that do not exist in ordinary spelling.

Therefore:

surface statistics belong first to the visible representation, not automatically to the hidden language beneath it.

Mechanism Family 1: Natural Language in an Unfamiliar Script

This is the most intuitive explanation.

Somebody wrote a real language using unfamiliar signs.

The attraction is obvious.

Voynichese has:

  • repeated word-like units;
  • frequency structure;
  • local vocabulary;
  • position-sensitive forms;
  • Currier variation;
  • syntax-like neighbour preferences;
  • labels and running text that are related but not identical.

Those are all things meaningful linguistic systems can produce.

Claire Bowern and Luke Lindemann’s 2021 review argued that linguistic methods can reveal language-like structure even without decipherment and gave reasons for taking a natural-language substrate seriously.

But the model still has to explain the strange surface.

Why are many glyphs so positionally constrained?

Why are word families so tightly related?

Why do certain prefixes and suffix-like sequences dominate?

Why is conditional predictability unusually strong?

A natural-language explanation therefore often becomes a natural-language-plus-transformation explanation.

Mechanism Family 2: Medieval Abbreviation and Scribal Compression

Medieval scribes compressed language aggressively.

Common endings could be replaced by signs.

Entire syllables could collapse into a mark.

Suspensions, contractions and superscript conventions reduced writing effort and conserved parchment.

Voynich glyphs such as minim-like strokes, benches, gallows and rare ligature-like forms therefore invite abbreviation comparison.

This family has one major advantage:

it is historically ordinary.

We do not need a futuristic machine.

But historical plausibility is only the first gate.

A serious abbreviation model must explain the whole system:

  • which glyphs are full letters;
  • which are abbreviations;
  • which combinations are ligatures;
  • why spaces occur where they do;
  • why Currier A/B differ;
  • why line positions matter;
  • why labels form a related population.

The article Abbreviation or Alphabet? owns this lane in detail.

Mechanism Family 3: Ciphertext

The manuscript has been called a cipher manuscript for generations.

That catalogue tradition should not be confused with a proven encryption method.

Classical simple substitution has long struggled with Voynich because the visible statistics do not behave like an ordinary one-symbol-for-one-letter substitution of common European plaintext.

But “cipher” is a much larger mechanism family than simple substitution.

  • homophonic substitution;
  • verbose substitution;
  • nulls;
  • multiple tables;
  • position-sensitive rules;
  • abbreviation plus ciphering;
  • codebook-like substitutions;
  • syllabic mappings.

What the 2025 Naibbe Result Actually Changes

Michael Greshko’s 2025 Cryptologia paper is important not because it claims to have deciphered Voynich.

It does not.

Its contribution is methodological.

Greshko designed a historically plausible hand-operable verbose homophonic substitution cipher—the Naibbe cipher—that can encrypt Latin and Italian plaintexts into ciphertexts displaying many Voynich-like statistical properties while remaining decryptable.

That changes one claim:

Voynich-like statistical behaviour is impossible to generate from a practical early-fifteenth-century cipher.

That strong claim is harder to maintain.

But the result does not establish:

  • that the Voynich Manuscript uses Naibbe;
  • that its plaintext is Latin;
  • that its plaintext is Italian;
  • that every Voynich feature is reproduced by the cipher;
  • that the manuscript is semantically solved.

In fact, the paper documents important mismatches and limitations.

This is exactly what makes it useful science rather than a “solution announcement”.

Naibbe demonstrates compatibility with a cipher family. Compatibility is not identity.

Mechanism Family 4: Generated or Meaningless Text

The most provocative alternative is that the text does not encode ordinary semantic prose at all.

A generator can produce local regularity.

It can reuse templates.

Modify neighbouring forms.

Create constrained syllable inventories.

Produce Zipf-like-looking distributions under some conditions.

Gordon Rugg famously proposed Cardan-grille-style generation as a historically plausible way to create large quantities of Voynich-like pseudo-text.

Later work has explored self-copying and other generative mechanisms.

The great lesson from these models is not that Voynich is proven gibberish.

It is that non-randomness is not enough.

A system can be structured without carrying the kind of meaning we expect from prose.

The Hard Part for Generated-Text Models Is Not One Word

Low-level imitation is only the beginning.

A manuscript-wide generator must explain why:

  • local vocabulary varies by page;
  • Currier A/B clusters exist;
  • labels differ from running text while remaining related;
  • certain forms prefer line beginnings or endings;
  • visual regimes correlate with textual behaviour;
  • multiple scribal hands or hand-like variations participate coherently;
  • some word families recur across distant folios;
  • the physical production process remains organised.

A generator that reproduces token shape but not manuscript-scale organisation explains only one layer.

Mechanism Family 5: Constructed Language or Specialised Notation

A system need not be an ordinary natural language or a cipher of one.

It could be a constructed language.

A private professional notation.

A mnemonic system.

A hybrid of words, abbreviations and technical signs.

This family is attractive because it can explain why Voynichese looks organised yet resists ordinary language identification.

It is also dangerously flexible.

If “constructed notation” can mean anything whenever a pattern fails, the hypothesis becomes unfalsifiable.

A useful constructed-system proposal must specify:

  • what the units are;
  • how they combine;
  • what constraints follow;
  • how the system is learned or used;
  • what new pattern the model predicts.

Mechanism Family 6: Hybrid Systems

This is perhaps the most historically ordinary possibility and the hardest to test.

Real manuscripts are messy.

A scribe may abbreviate language before enciphering it.

A cipher may preserve some spaces but not others.

Technical names may remain in plaintext while prose is encoded.

Labels may use a different convention from running text.

One section may inherit a source tradition with different abbreviations.

Several scribes may apply the same rules differently.

A hybrid model can therefore explain heterogeneity naturally.

It also creates too many degrees of freedom if unconstrained.

The rule is the same:

every extra mechanism should buy explanatory power somewhere else.

What Each Family Explains Naturally

Natural language

Explains document-scale semantic organisation, vocabulary, syntax-like behaviour and communicative purpose naturally.

Abbreviation

Explains unusual compact forms, ligatures, position-sensitive marks and medieval-looking compression naturally.

Cipher

Explains unreadability, transformed statistics and deliberate opacity naturally.

Generated text

Explains strong local regularity, near-neighbour word families and some low-level statistical anomalies naturally.

Constructed notation

Explains an unfamiliar but internally coherent representation system naturally.

Hybrid

Explains why several apparently incompatible properties can coexist.

What Each Family Has to Pay For

No explanation gets the manuscript for free.

  • Natural language: must explain the unusually constrained surface.
  • Abbreviation: must reconstruct a consistent expansion system.
  • Cipher: must reproduce the observed manuscript-wide behaviour and decrypt coherently.
  • Generated text: must explain higher-level organisation, not only low-level mimicry.
  • Constructed notation: must specify actual rules rather than naming the unknown.
  • Hybrid: must avoid becoming an unlimited patchwork of exceptions.

Why Statistical Fit Does Not Equal Mechanism Identity

Suppose two mechanisms generate the same entropy.

They are not therefore the same mechanism.

Suppose one cipher reproduces word-length distribution and positional constraints.

That is important.

It still must explain page vocabulary, Currier variation, labels, images, line effects and exceptions.

Likewise, a generated-text model that reproduces local repetition has not automatically explained meaninglessness.

Statistical fit is evidence that a mechanism can produce a property. It is not proof that the historical manuscript used that mechanism.

Why “Language-Like” Does Not Mean “Language Identified”

Voynichese can look language-like in several statistical senses.

That does not identify Latin, Italian, German, Hebrew, Nahuatl, Pahlavi or any other proposed language.

Language identification requires discriminating evidence.

Shared statistical laws are not enough because many languages and non-language generators can share them.

The stronger evidence would be:

  • consistent phonological mapping;
  • repeatable morphology;
  • syntax that predicts unseen cases;
  • stable vocabulary linked to independent contexts;
  • coherent translation across long passages;
  • exceptions explained by one mechanism rather than ad hoc repairs.

Why “Gibberish-Like” Does Not Mean “Meaningless”

The same asymmetry applies in reverse.

A cipher can make meaningful plaintext look highly artificial.

An abbreviation system can destroy ordinary word shapes.

A specialist notation can be opaque to outsiders.

So no one low-level oddity is sufficient to conclude that the text is meaningless.

The Manuscript-Scale Test

The decisive mechanism should explain more than tokens.

  • Why do visual sections differ?
  • Why do Currier A and B exist?
  • Why do labels form a distinct population?
  • Why do line beginnings and endings matter?
  • Why do some word families remain local?
  • Why are there multiple scribal hands or hand-like variations?
  • Why does the same system survive across herbal, zodiac, balneological and starred material?
  • Why do rare glyphs behave the way they do?
  • How does the mechanism cope with missing leaves and uncertain ordering?
  • What happens on unseen text?

This is the burden described in What a Real Voynich Decipherment Must Survive.

Primary School: Secret Code or Made-Up Pattern?

Write two strings on the board.

One is a secret substitution code for a real sentence.

The other is generated by a simple pattern rule.

Both can look strange.

Ask the child:

Can you tell which one has a hidden message just by looking?

Usually not.

The child has just discovered the Voynich mechanism problem.

Secondary School: Compare Explanatory Costs

Make six columns:

  • natural language;
  • abbreviation;
  • cipher;
  • generated text;
  • constructed notation;
  • hybrid.

For each observed feature, mark which mechanism explains it naturally and which needs an extra assumption.

Students quickly learn that the best theory is not always the one explaining the most famous clue.

It is the one paying the lowest total explanatory cost across many clues.

Reader Checklist: Before You Declare Voynich a Cipher, Language or Hoax

  1. What exact mechanism is proposed?
  2. Which observed properties does it reproduce?
  3. Which does it fail?
  4. Was the mechanism designed after seeing those properties?
  5. Does it work on unseen passages?
  6. Can it be performed with historically plausible tools?
  7. Does it preserve or explain Currier A/B?
  8. Does it explain labels and running text?
  9. Does it account for line-position effects?
  10. Does it explain visual-text relationships?
  11. Does it produce coherent long-form output?
  12. Are exceptions governed by one rule or patched individually?
  13. Could a competing mechanism produce the same statistics?
  14. What observation would falsify the proposal?

Frequently Asked Questions

Is Voynich definitely a cipher?

No. Ciphertext is one viable mechanism family. No specific cipher has been established as the historical mechanism.

Does Naibbe solve Voynich?

No. It demonstrates that a historically plausible hand-operable cipher can reproduce many Voynich-like properties while encoding real Latin or Italian plaintext. That keeps the cipher hypothesis viable; it does not identify the manuscript’s key or plaintext.

Does language-like structure prove natural language?

No. Generated systems and ciphers can reproduce some language-like statistics. Natural language remains plausible, but its mechanism must explain the complete surface system.

Does unusual repetition prove gibberish?

No. Repetition can arise from generation, morphology, abbreviation, cipher structure or specialised notation.

Could it be several mechanisms at once?

Yes. Hybrid systems are historically plausible. But every additional mechanism must explain independent evidence rather than merely rescue failures.

Research Foundations

The Final Idea

Voynichese does not need one more adjective.

Language-like.

Cipher-like.

Generated-looking.

Abbreviated.

Artificial.

Those words are beginnings.

What we need is a mechanism that reproduces the manuscript’s behaviour and then makes new predictions.

The winning explanation will not be the theory that Voynich resembles. It will be the mechanism the manuscript cannot escape.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading