Voynich is an optimiser’s dream.
Thousands of unknown tokens.
Many possible glyph values.
Many candidate languages.
Many ways to segment sounds.
Give a computer enough freedom and one instruction—find me readable words—and it will work very hard to obey.
That is exactly why optimisation can become dangerous.
In a 2026 reproducible open-analysis project, researchers let a computer search for symbol-to-sound assignments that would make Voynich tokens resemble words in Latin, Greek or Turkic lexicons. The search happily produced apparent dictionary matches across more than half the manuscript.
Then the researchers performed the control that matters.
They scrambled the dictionaries into nonsense while preserving the search opportunity.
The optimiser still found convincing-looking “translations”.
The project later reports null hit rates around 57–74% under search-matched scrambled-lexicon controls.
That is not evidence that Voynich means nothing.
It is evidence that dictionary-hit optimisation can make nonsense look meaningful at alarming scale.
If the same search finds “translations” in real dictionaries and scrambled nonsense dictionaries, the search has measured its own flexibility more than it has measured Voynich meaning.
Quick Read
- The Voynich Project is a 2026 reproducible open-analysis release maintained by an independent researcher; it is not peer-reviewed consensus.
- One experiment optimised symbol-to-sound mappings against Latin, Greek and Turkic dictionaries.
- The optimiser reportedly produced apparent matches for more than half the manuscript.
- The same procedure was then run against scrambled nonsense dictionaries.
- It performed similarly, showing that high dictionary-hit rates can arise from search flexibility alone.
- The project reports search-matched null hit rates around 57–74% in its later audit.
- This does not prove every published translation is false by itself.
- It proves that dictionary-hit percentage is weak evidence unless compared with a search-matched null.
- A real translation must outperform scrambled controls, preserve one fixed mapping, generalise to held-out text, respect manuscript structure, and recover external meaning rather than merely dictionary resemblance.
- The experiment is one of the clearest demonstrations of why Voynich “translation” requires falsification, not just optimisation.
Why Search Finds What It Is Asked to Find
Optimisation is powerful because it explores many possibilities efficiently.
Suppose every Voynich glyph can receive several possible sound values.
Suppose common clusters can be merged or split.
Suppose the target lexicon contains thousands of words.
The search space becomes enormous.
Among enough possibilities, some mapping will make many strings resemble real words.
This is not dishonesty.
It is multiple testing expressed as an algorithm.
The Wrong Question
A weak translation search asks:
Can I find a mapping that makes many Voynich tokens look like dictionary words?
The answer can be yes even when the target dictionary is fake.
The stronger question is:
Does the best real-language mapping outperform equally searchable nonsense controls by enough margin to survive held-out testing?
Why Scrambled Dictionaries Are Such a Strong Control
A fair null should preserve the opportunities available to the optimiser while removing the claimed semantic truth.
If real dictionaries contain words of certain lengths and phonotactic shapes, a crude random-character dictionary may be too easy to reject.
A scrambled lexicon can preserve much of the search geometry while destroying genuine lexical identity.
If the optimiser succeeds similarly on both, the apparent success rate belongs mainly to the optimisation machinery.
This is exactly the kind of null-model discipline Voynich research needs.
Why “More Than Half the Book” Sounds Stronger Than It Is
Coverage percentages feel intuitive.
Sixty percent translated.
Seventy percent matched.
Eighty percent readable.
But a percentage is only meaningful relative to a null.
If a nonsense dictionary yields 65% under the same search and the real dictionary yields 67%, then 67% is not evidence of translation.
The search baseline has already consumed almost the entire apparent effect.
This is the core lesson of the reported 57–74% null range.
The Optimiser Can Exploit Repeated Voynich Structure
Voynich is highly constrained.
Many tokens belong to families.
Some components recur predictably.
This helps an optimiser because one clever mapping can create many related dictionary-like hits at once.
That can make the output look linguistically coherent even if the mapping is only exploiting regular surface families.
Structural regularity makes the false translation more convincing.
It does not make it true.
Why Human Solvers Face the Same Trap
Humans optimise too.
We simply do it less explicitly.
A solver proposes a sound value.
Changes another.
Moves a space.
Allows one abbreviation.
Then notices a familiar word.
That word encourages another adjustment.
Soon a coherent reading appears.
The process can be sincere and still be overfitted.
The Patch Problem owns that rescue-freedom issue.
Real Language Must Win Against the Same Search Budget
One common mistake is comparing:
- a highly optimised real-language mapping;
- a weak random null with little optimisation.
That is unfair.
The null must receive the same search budget.
Same number of mappings tried.
Same flexibility.
Same objective.
Same stopping rule.
Only then does the difference between real and scrambled targets mean anything.
A Training/Test Split Is Essential
Even if a real-language lexicon beats scrambled controls on the training corpus, the mapping must be frozen before looking at held-out Voynich text.
Then ask:
- Does the same glyph mapping still produce valid forms?
- Does coverage remain above the null?
- Do frequent words remain stable?
- Do section differences make linguistic sense?
- Does the mapping explain labels and prose without separate rulebooks?
If the mapping collapses on unseen pages, the optimiser learned the training sample rather than the manuscript mechanism.
Dictionary Match Is Not Sentence Meaning
A token resembling a dictionary entry is only the first layer.
A real reading must create coherent relationships among tokens.
Grammar.
Function words.
Agreement.
Document structure.
Semantic references to images or external knowledge.
This is why the existing Semantic-Grounding Problem remains distinct from the Translation Search Trap.
One asks how meaning becomes externally anchored.
This article asks whether the search procedure can create false lexical evidence before semantics even enters.
Why Latin, Greek and Turkic Are Good Demonstrations
These language families differ substantially.
If one flexible search can make Voynich look impressively compatible with several unrelated targets, that is itself a warning.
The manuscript cannot simultaneously be ordinary Latin, Greek and Turkic under mutually incompatible sound systems.
The shared success therefore tells us more about the optimiser’s degrees of freedom than about language identity.
Why Turkic Gets an Independent Test
The same project separately built vowel-harmony detectors calibrated on Turkish and Finnish and reports no corresponding Voynich harmony at sign or learned-unit level.
It also built a period-oriented Cuman corpus for a fairer medieval Turkic comparison.
This illustrates an important principle.
Dictionary matching is weak.
Language-specific structural predictions are stronger.
A proposed language should be asked to bring its grammar, phonotactics and typology—not merely its word list.
The Search Trap Explains Why Voynich Headlines Keep Happening
The manuscript has enough regularity to support beautiful mappings.
Enough ambiguity to permit alternatives.
Enough short tokens to match many dictionary entries.
Enough illustrations to provide suggestive context.
That combination is perfect for overfitting.
A solver can produce something readable and experience genuine conviction.
The scientific response is not mockery.
It is a stronger control.
What a Real Translation Search Should Require
- Freeze the transcription.
- Freeze segmentation rules.
- Declare all allowed glyph mappings.
- Declare the target lexicon before search.
- Construct search-matched scrambled or pseudo-lexicon controls.
- Use the same optimisation budget on real and fake targets.
- Train on one subset of Voynich.
- Freeze the mapping.
- Evaluate on held-out pages.
- Test language-specific grammar and phonology.
- Require external semantic anchors.
- Publish null results and failed languages.
The Best Evidence Is Something the Optimiser Cannot Choose
A secure external crib is powerful because the optimiser does not get to choose the answer.
A known month name.
A secure plant label.
A bilingual annotation.
A matching historical source text.
The target exists independently.
That is why the Crib Problem remains central.
Blind Prediction Is the Semantic Version of the Same Rule
The Blind Prediction Problem applies the same logic to images.
Do not let the theory see the target first.
Prediction before reveal prevents optimisation from quietly absorbing the answer.
What the Translation Search Trap Does Not Prove
- It does not prove Voynich is meaningless.
- It does not prove no language can ever be identified.
- It does not prove every historical translation claim is equally bad.
- It does not prove Latin, Greek or Turkic are impossible under every imaginable transformed system.
- It does not replace specialist philology with one optimisation experiment.
- It does prove that dictionary-hit rate alone can be badly misleading.
What It Can Give Us
- A quantified measure of false translation plausibility.
- A search-matched null standard.
- A reason to distrust raw dictionary coverage.
- A clean separation between lexical resemblance and language-specific structure.
- A practical falsification gate for human and AI translation systems alike.
Reader Checklist
- How many mappings were searched?
- How flexible was segmentation?
- Was the dictionary chosen in advance?
- What scrambled or fake dictionary control was used?
- Did the null receive the same search budget?
- What is the real-minus-null effect size?
- Was the mapping frozen before held-out testing?
- Does the target language’s grammar fit?
- Does the target language’s phonology fit?
- Are external semantic anchors present?
- Could the optimiser produce similar readability from nonsense?
Research Foundations
- The Voynich Project — reproducible open-science investigation (2026 public research release; not peer reviewed).
- The Blind Prediction Problem.
- The Patch Problem.
- The Crib Problem.
The Final Idea
The most dangerous Voynich translation is not necessarily the one that looks ridiculous.
It is the one produced by a powerful search that has been allowed enough freedom to make almost anything look convincing.
A real translation should survive the moment we give the optimiser an equally attractive nonsense world and ask it to choose correctly.