VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Codebook Problem: What If Voynich Tokens Are Codes Rather Than Spelled Words?

What if a Voynich “word” is not a word at all?

Not an encrypted sequence of letters.

Not a compressed spelling.

Not a grammatical bundle.

A code.

One visible token stands for an entire plaintext unit.

A person.

A place.

A plant.

An ingredient.

A common word.

A syllable.

A morpheme.

A concept.

This is not an ahistorical fantasy.

European cryptography from the late medieval and early modern periods used nomenclators: hybrid systems combining ordinary letter substitution with a special list of codes for larger plaintext units.

Early nomenclatures frequently focused on people and places.

Over time, surviving keys expanded toward broader vocabularies and more complex codebooks.

Some systems combined nomenclature with homophonic substitution.

Nulls.

Codes for n-grams.

Several cryptographic resolutions could coexist inside one message.

That immediately gives the Voynich debate a powerful possibility.

Perhaps some Voynich units should not be deciphered character by character because the intended reader looked them up as wholes.

But the manuscript creates an immediate problem for a simple codebook story.

Voynich tokens are not arbitrary serial numbers.

They have severe internal structure.

Related-looking families.

Positional glyph rules.

Currier differences.

Boundary effects.

If each token is a whole-word code, why should the codes themselves look so linguistically or mechanically related?

That is the Codebook Problem.


Quick Read

Direct answer: a nomenclator or mixed codebook is historically plausible for the Voynich period and could hide names, words, syllables or concepts behind dedicated ciphertext units. However, a simple arbitrary word-code list does not naturally explain Voynich’s dense internal token families, character-position constraints, graded boundaries and regime differences. A viable codebook theory must specify the encoded unit, explain how codes are generated and retrieved, recover a compact key or key-generating rule, and predict unseen labels and prose without turning every rare token into an unconstrained dictionary entry.

  • Nomenclators were real hybrid encryption systems used from the late medieval period onward.
  • They commonly combined alphabetic substitution with dedicated codes for larger plaintext entities.
  • Early surviving nomenclatures often encoded names of people and places.
  • Later systems expanded into larger dictionaries covering a wider range of linguistic material.
  • Historical nomenclators could also include homophones, n-gram substitutions and nulls.
  • Therefore one ciphertext message could contain several kinds of units at once.
  • A Voynich token could in principle represent one plaintext word, name, syllable, morpheme or concept.
  • Labels are an obvious testing ground because names and identifiers are historically natural nomenclator targets.
  • Quire 20 record-like units are another plausible place to test codebook models.
  • A purely arbitrary codebook should not automatically generate strong internal glyph-position grammar.
  • Voynich word families therefore require either structured code assignment, compositional codes or another layer beyond simple lookup.
  • The manuscript’s large singleton-rich vocabulary makes a one-code-per-word model potentially expensive.
  • A missing codebook is historically possible but cannot be assumed to contain whatever mapping a modern theory needs.
  • f57v and other key-like sequences remain structural evidence, not accepted recovered nomenclators.
  • A codebook solution should become increasingly routine as the key fills in; unseen tokens should decline, not remain permanently miraculous.
  • No accepted Voynich nomenclator or codebook has been recovered.

What a Nomenclator Is

A nomenclator is not merely a substitution alphabet.

It is a hybrid.

Ordinary letters may receive ciphertext substitutes.

Then a separate list assigns special codes to high-value plaintext items.

For example:

  • the Pope → one dedicated symbol or number;
  • Venice → another;
  • the Emperor → another;
  • a frequently used diplomatic term → another.

The code bypasses spelling.

That is the central idea.

Once a plaintext entity has a code, the ciphertext does not need to reveal its letters individually.

The Fifteenth Century Makes the Idea Historically Serious

Historical cryptology research places nomenclator systems in use from the fourteenth and fifteenth centuries onward, especially in Italian diplomatic culture.

Surviving fifteenth-century nomenclature lists tend to be smaller and thematically organised compared with the large later codebooks.

People.

Places.

Political entities.

As systems developed, the nomenclator could absorb more common words, linguistic fragments and specialised vocabulary.

This matters because Voynich’s radiocarbon horizon overlaps a period in which mixed-granularity cipher design is historically plausible.

It does not mean a medical or botanical book would normally use a diplomatic nomenclator.

Historical possibility and manuscript function remain separate questions.

Codebook Is Not the Same as Homophony

The Homophony Problem asks:

can one plaintext unit have several ciphertext disguises?

The Codebook Problem asks:

what size plaintext unit does one ciphertext object represent?

A codebook can be one-to-one.

One plaintext word.

One code.

A homophonic codebook can be many-to-one.

One plaintext word.

Several codes.

The mechanisms can coexist but should not be conflated.

Granularity and multiplicity are different axes.

Why Labels Are the Obvious Place to Look

Historical nomenclators were especially useful for names.

Names are long.

Politically important.

Often predictable.

Give each one a dedicated code and the intercepted message loses a valuable crib.

Voynich labels are therefore natural codebook candidates.

A star label could be one code.

A plant label one code.

A figure identifier one code.

This has one attractive consequence.

The label need not resemble the historical spelling of the underlying name at all.

That explains why proper-name cribs may fail.

But it creates another burden.

Without the codebook, how do we distinguish one arbitrary label code from another?

A Pure Arbitrary Codebook Predicts Weak Internal Morphology

Suppose each common plant receives a random five-symbol code.

Rose = AKRDP.

Sage = LMUQT.

Thyme = ZBCVN.

The internal letters of those codes need not obey root-and-affix relationships.

They can be arbitrary identifiers.

Voynich is different.

Its visible tokens occupy a highly constrained internal space.

Some glyphs strongly prefer beginnings.

Others endings.

Words form dense near-neighbour families.

A random codebook does not explain this naturally.

The code assignments themselves must be structured.

A Structured Codebook Is More Interesting

Codes do not have to be random.

A codebook designer can group related entities.

Plant family A receives codes beginning with X.

Preparation class B receives endings Y.

Quantity or state occupies another slot.

Now related codes look like word families.

The token itself becomes a compact structured record.

This begins to resemble a nomenclature, classification code or compositional notation more than an arbitrary diplomatic codebook.

It is plausible.

It is also close to inventing a semantic ontology unless the slots are discovered from independent distributional evidence.

Codebook and Compositional Code Are Different

A pure codebook says:

XYZ = “mandrake”.

The internal X, Y and Z need not mean anything separately.

A compositional code says:

X = plant class, Y = root, Z = preparation state.

Now each position carries information.

Voynich’s positional character grammar makes compositional codes especially tempting.

But the burden is stronger.

Each slot should correlate with independent page, image, label or sequence variables.

Otherwise the slot meanings are assigned after the fact.

Whole-Word Codes Can Explain Weak Exact Token Syntax

If each token is a code for a lexical or conceptual unit, exact token-to-token transitions can be sparse.

A technical text may contain hundreds of rare entities.

Two exact codes need not repeat in the same sequence often.

Yet category-level order can remain strong.

Ingredient code → operation code → quantity code.

This resembles the distinction already visible in Voynich between weak exact-token sequence and stronger class/edge structure.

A codebook model therefore should not be tested only at exact token identity.

It should recover latent code classes.

The Open Vocabulary Is Both a Strength and a Problem

Voynich has a large long tail of rare and singleton forms.

A codebook can explain this naturally if the underlying domain contains many unique names or entities.

One plant.

One star.

One recipe.

One person.

Each receives one code.

But the same explanation can become too easy.

Every unexplained singleton becomes “a unique codebook entry”.

The model never fails because the missing dictionary can contain anything.

A valid codebook theory therefore needs constraints from code structure, document role or external referents.

Currier A and B Could Be Different Codebooks—or Different Coding Dialects

Perhaps Currier A and B use different sections of one nomenclator.

Or different code tables.

Or the same categories encoded with different conventions.

This can explain surface distribution shifts without requiring two spoken languages.

Again, the model makes predictions.

Equivalent visual or document roles should map to related source categories across regimes.

If A and B codebooks are completely unrelated, key complexity doubles.

Historical plausibility is not enough; economy matters.

The Missing Key Is a Historical Possibility, Not an Explanation

Codebooks can be lost.

Cipher keys were often separate operational documents.

A manuscript can survive after its decoding apparatus disappears.

This is genuinely relevant to Voynich.

But “the key was lost” cannot answer:

  • why q behaves the way it does;
  • why word lengths are constrained;
  • why paragraph starts are special;
  • why edge glyphs couple;
  • why labels differ from prose.

The lost key can explain our ignorance.

It cannot replace a model of the surviving ciphertext.

f57v Is Not Yet the Missing Codebook

f57v looks key-like.

It contains a repeated approximately seventeen-position sign sequence.

That makes it an extraordinary structural test.

But it does not give us readable plaintext values paired with those signs.

A codebook needs mappings.

f57v currently provides mysterious structure rather than a recovered dictionary.

The Paratext Problem Becomes More Important Under a Codebook Model

A long technical codex using dedicated codes needs retrieval.

How does the reader know which code family applies?

Where does one table end?

How are related entries found?

Where is the key?

Perhaps trained users memorised the system.

Perhaps labels and section structure were sufficient.

Perhaps the external key is lost.

All are possible.

But a codebook model increases rather than decreases the importance of the manuscript’s missing navigational layer.

Quire 20 May Be More Code-Like Than Narrative Prose

The starred pages are often described as record-like.

Short entries.

Repeated markers.

Formulaic structure.

This kind of document can naturally use codes.

A code may represent a substance, procedure, diagnosis or cross-reference.

Yet nothing in the current evidence identifies one star entry as a decoded codebook record.

Document compatibility should generate tests, not conclusions.

A Medical Codebook Is Plausible but Easy to Romanticise

Medicine contains repeated named entities.

Plants.

Diseases.

Preparations.

Measures.

Operations.

A specialised notation could encode them compactly.

But the visual resemblance of Voynich to medical books cannot define the code values.

Otherwise the images become the dictionary that validates itself.

The same firewall from the Crib Problem applies.

A Nomenclator Can Mix Spelled Text and Codes

This may be the most relevant historical lesson.

Not every ciphertext unit needs the same granularity.

One ordinary word may be encrypted letter by letter.

A proper name may receive one whole code.

A common syllable may have its own sign.

A null may be inserted between them.

This hybrid architecture could explain why some Voynich phenomena look letter-like while others look token-like.

It also creates a difficult inference problem.

Which unit is operating where?

A successful model must discover the switching rule.

Mixed Granularity Can Explain Why Cribs Fail

Suppose a plant name receives a whole-word code.

Trying to align its letters with Voynich glyphs will fail.

Suppose the surrounding prose is letter-by-letter ciphertext.

The same character sequence can then behave differently depending on unit type.

This makes a mixed nomenclator powerful enough to explain many failed phonetic cribs.

But power is not proof.

The switching points must be recoverable from structure rather than chosen whenever one reading fails.

Codebook Size Must Be Historically and Operationally Plausible

A codebook with ten entries is easy to memorise.

A codebook with thousands is not.

Large historical nomenclatures required organised keys.

Voynich contains thousands of distinct visible token types under common tokenisations.

If every type receives a unique unrelated plaintext entry, the implied codebook becomes enormous.

That may be historically later than the simplest fifteenth-century nomenclatures and operationally cumbersome unless the codes are generated compositionally.

Therefore key size itself can discriminate theories.

Word Families Can Compress the Implied Codebook

Suppose related visible tokens are not separate arbitrary codes.

They are variants generated from one base code plus modifiers.

Now the key becomes smaller.

Base code = plant.

Ending = preparation state.

Prefix = quantity class.

This is attractive because it explains family geometry and reduces memorisation.

But it is no longer a simple codebook.

It is a compositional grammar.

The modifiers must now have stable cross-base effects.

The Edge Problem Tests Mixed-Granularity Codes

Recent work reports strong relationships across token edges despite weak exact token-to-token predictability.

A codebook model can explain weak exact token order if codes are sparse.

It must still explain why code endings and beginnings interact.

Possibilities include:

  • grammatical markers outside the code itself;
  • code-class compatibility;
  • state transfer between groups;
  • transcription splitting one larger code unit.

This is a useful new discriminator.

An arbitrary lookup list has no reason for adjacent code edges to couple unless the code design explicitly builds that relation.

The Key-Like Sequences Must Become Useful Under the Theory

Voynich contains several unusual sequences that look more table-like or key-like than ordinary prose.

f57v is the most famous.

A mature codebook theory should eventually explain whether these pages are:

  • alphabets;
  • code classes;
  • calibration tables;
  • unrelated diagrams.

The theory should not need them to be keys merely because that would be convenient.

But once a code architecture is independently recovered, key-like pages become powerful validation cases.

A Codebook Theory Must Generate a Retrieval Story

How did the intended user decode the book?

Look up every token in a separate table?

Memorise common classes?

Use visual context to choose a code family?

Decode only labels while prose used another mechanism?

A historical system is an operating procedure, not just a mapping.

The user experience matters.

A codebook so complex that its intended reader could not plausibly use it inside a working manuscript is a poor historical fit.

The Almost-Correctionless Manuscript Tests Retrieval Fluency

If scribes were encoding through a codebook, they either knew the system well or consulted a key efficiently.

The surviving manuscript has few obvious corrections.

This is compatible with trained operation.

It is less compatible with an enormous unfamiliar lookup system being improvised during writing.

A codebook mechanism must fit not only statistics but human performance.

A Codebook Theory Should Eventually Discover Repeated Plaintext Categories

Even if exact plaintext values remain unknown, code classes can be tested.

Codes near plants may cluster into one family.

Codes at paragraph beginnings another.

Codes attached to stars another.

If the same latent classes recur across visual contexts predictively, the codebook hypothesis gains structure.

If every token remains isolated, the theory becomes a missing dictionary with no recoverable grammar.

What Survives the Codebook Work

  • Nomenclators and mixed-granularity codes are historically real and relevant to the broader Voynich period.
  • They can combine letter substitution with dedicated codes for names, places and other larger plaintext units.
  • Some historical systems also combine homophones, n-grams and nulls.
  • Voynich labels are legitimate places to test codebook hypotheses.
  • Record-like text such as Quire 20 is also structurally compatible with coded entries.
  • A simple arbitrary codebook does not naturally explain Voynich’s internal token grammar.
  • Structured or compositional codes can explain more but create stronger cross-token predictions.
  • The large singleton-rich vocabulary makes unconstrained one-code-per-token theories expensive.
  • The absence of a surviving key is historically possible but cannot substitute for evidence in the ciphertext.
  • No accepted Voynich codebook or nomenclator has been recovered.

What Does Not Survive as Established Knowledge

  • Voynich tokens are proven whole-word codes.
  • Voynich labels are proven nomenclator entries.
  • f57v is a recovered codebook.
  • The missing leaves contained the lost key.
  • Every singleton token represents one unique object.
  • Currier A/B are proven different nomenclators.
  • A medical-looking page proves a medical code ontology.
  • The existence of historical nomenclators proves Voynich is ciphered.

Mixed-granularity coding survives as a mechanism class.

The Voynich dictionary does not.

A Better Codebook Analysis

  1. Specify the hypothesised plaintext unit: name, word, syllable, morpheme or concept.
  2. Separate arbitrary lookup codes from compositional codes.
  3. Estimate the implied key size.
  4. Compare with historically appropriate nomenclators rather than modern abstract codebooks.
  5. Test labels and running text separately.
  6. Ask whether word-family components have stable modifier effects.
  7. Test Currier and hand transfer.
  8. Explain how code classes interact across token edges.
  9. Recover a plausible reader lookup procedure.
  10. Freeze the inferred key or code-generation grammar.
  11. Predict unseen entries.
  12. Do not assign every failure to an unpreserved dictionary entry.

What Would Count as a Real Codebook Breakthrough?

Imagine a set of label tokens is analysed without using the pictures to assign meanings.

The tokens fall into compact structural classes.

One class predicts astronomical labels.

Another plant-component labels.

A small number of internal modifiers recur across otherwise unique codes and correlate with independently measurable properties.

A historical nomenclator-like architecture is frozen.

On unseen pages, the code classes predict which visual/document role a token belongs to before the image is consulted.

Then one independent crib fixes several actual values.

Those values propagate through the code grammar and begin recovering other entries without adding arbitrary dictionary mappings.

That would be a codebook breakthrough.

The missing Voynich codebook becomes scientific evidence only when enough of its organisation can be reconstructed from the manuscript that the unseen entries begin predicting themselves.

Primary School: One Symbol for One Whole Word

Invent a symbol meaning “Singapore”.

Another meaning “school”.

Now a child can write two complicated words without spelling either one.

The exercise shows why reading every ciphertext sign as one letter can fail if the sign is really a whole-word code.

Secondary School: Mixed Granularity

Encode an ordinary sentence letter by letter, except replace every person and place name with one dedicated code.

Now ask a student to infer the system.

Some ciphertext groups behave alphabetically.

Others do not.

The same message uses two resolutions.

JC and Adult Readers: Latent Granularity

At a higher level, the codebook problem is a latent-granularity problem.

Observed units may correspond to source letters, n-grams, morphemes, words or concepts.

A flexible model can switch among these arbitrarily and fit almost anything.

Therefore the switching rule itself must be inferred from reproducible distributions and penalised for complexity.

The best model is the smallest mixed-resolution grammar that predicts which source granularity operates where.

A Parent and Teacher Guide

  1. Understand that ciphertext need not be letter-by-letter.
  2. Separate code granularity from homophonic multiplicity.
  3. Use historical nomenclators as real controls.
  4. Do not turn every rare token into a free dictionary entry.
  5. Demand structure inside the codebook.
  6. Estimate the implied key size.
  7. Ask how the intended reader would retrieve values.
  8. Make the inferred organisation predict unseen entries.

The wider lesson is:

before decoding an unknown symbol, first establish what size thing the symbol is supposed to represent.

Reader Checklist: Before You Call Voynich a Codebook Cipher

  1. What plaintext unit does one code represent?
  2. Are codes arbitrary or compositional?
  3. How large is the implied dictionary?
  4. Is that size historically plausible?
  5. How would the intended reader retrieve values?
  6. Are labels being analysed separately?
  7. Do internal code components have stable roles?
  8. Do code families predict document categories?
  9. Does the theory explain token-edge dependence?
  10. Does it explain Currier A/B?
  11. Does it require a different unseen dictionary for every section?
  12. Can an independent crib fix actual values?
  13. Does the fixed system predict unseen codes?
  14. Is “lost key” being used to protect the theory from every failure?

Frequently Asked Questions

What is a nomenclator?

A historical hybrid cipher system combining ordinary substitution with dedicated codes for larger plaintext units such as names, places, titles, words or other entries.

Were nomenclators used in the fifteenth century?

Yes. They are documented from late medieval/early Renaissance European cryptographic practice, including Italian diplomatic contexts.

Could one Voynich token mean one whole word?

In principle yes. No accepted mapping demonstrates that this is how Voynich tokens function.

Why is the internal token grammar a problem?

Random whole-word codes need not share strong prefix, suffix and positional structure. Voynich does, so a codebook theory needs structured or compositional code assignment rather than an arbitrary lookup list.

Could f57v be the codebook?

It is key-like and structurally important, but it lacks accepted readable values paired with its repeated signs. It is not a recovered nomenclator.

What is the strongest current conclusion?

Mixed-granularity coding is historically plausible and should remain in the Voynich mechanism space, but the manuscript’s internal token grammar and enormous visible vocabulary place strong constraints on any simple codebook explanation.

Related eduKateSG Reading

Research and Further Reading

The Final Idea

The codebook hypothesis is attractive because it lets us stop asking every Voynich sign to behave like a letter.

Perhaps one token means a whole plant.

A star.

A diagnosis.

A word.

A syllable.

History permits that kind of cryptographic granularity.

But Voynich refuses to become a pile of arbitrary serial numbers.

Its codes, if they are codes, have architecture.

They resemble one another.

They know where they are inside a token.

Their edges relate.

Their frequencies drift by regime.

If Voynich is a codebook, the breakthrough will not come from imagining the missing dictionary. It will come from reconstructing enough of the dictionary’s internal organisation that the manuscript itself starts telling us what an unseen code is allowed to mean.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading