Voynich research has spent a century arguing about whether one line can be translated.
In August 2026, Matthew Dominik deposited something much larger.
A 1,609-line English translation edition.
Not a word-by-word substitution alphabet.
Not a claimed recovery of historical pronunciation.
Not a complete dictionary.
The project describes its result as a controlled translation generated from a frozen compositional grammar with explicit unresolved material retained in brackets.
The deposit reports a 34,776-token working corpus across 4,084 lines, a frozen structural classifier that licenses 30,534 tokens under ZL3b normalization, and a fully constrained 1,609-line benchmark containing 12,495 token positions, of which 12,294 receive what the author calls a defensible functional-role interpretation.
Those are unusually specific claims.
They deserve unusually specific scrutiny.
The important question is not whether 1,609 English lines exist. They do. The important question is what chain of evidence makes those English lines uniquely compelled by the Voynich source rather than one coherent interpretation generated by a flexible grammar.
Quick Read
- The August 2026 work by Matthew Dominik is a Zenodo preprint, not accepted scholarly consensus.
- It presents a 1,609-line English controlled translation derived from a frozen compositional grammar.
- The full project corpus contains 34,776 normalized tokens across 4,084 lines.
- Under canonical ZL3b normalization, the frozen structural classifier reportedly licenses 30,534 tokens, or 87.802%.
- The constrained 1,609-line benchmark contains 12,495 source-token positions.
- The deposit reports 12,294 of those positions, or 98.3914%, receiving a defensible functional-role interpretation, with 201 positions retained as opaque.
- The author explicitly states that 98.3914% is not whole-manuscript lexical accuracy because benchmark membership itself depends on model constraint.
- Adversarial replication on RF1b reportedly preserves the principal role architecture but lowers whole-corpus coverage to 86.065%.
- The work does not claim recovery of historical phonology or a complete dictionary for every root.
- Northwest Caucasian / Abkhaz–Abaza is retained as a structural analogue rather than a recovered source language.
- The main scientific burden is therefore uniqueness: do the frozen grammatical roles force the English interpretation strongly enough that competing grammars cannot produce comparably coherent output?
- A controlled translation can be an important research object without yet being an accepted decipherment.
First: What Does “Controlled Translation” Mean Here?
The deposit uses the phrase deliberately.
The author argues that the output goes beyond a morphological hypothesis because the system produces a source-order English edition over a large constrained benchmark.
At the same time, the project explicitly refuses the stronger statement:
the Voynich Manuscript has been completely deciphered.
The proposed middle category is:
an English translation of a constrained corpus under a reproducible compositional grammar, with unresolved lexical material exposed rather than silently guessed.
This framing is intellectually useful because it separates several achievements that are often bundled together.
- structural parsing;
- functional-role assignment;
- lexical interpretation;
- sentence-level English rendering;
- historical-language identification;
- phonological recovery.
A system can succeed at some and remain uncertain at others.
Why 1,609 Lines Is a Serious Scale
Voynich solutions often collapse because they explain one page.
Or one plant.
Or one circular diagram.
A 1,609-line benchmark is different.
At that scale, local coincidence becomes harder to sustain.
Repeated grammatical roles must recur.
Exceptions accumulate.
Section differences appear.
The model must keep assigning structures under many contexts.
Scale is therefore a genuine strength.
But scale alone cannot solve non-uniqueness.
A flexible parser can apply the same interpretive grammar repeatedly over a large corpus and still be wrong about what the grammar represents historically.
The Difference Between Coverage and Accuracy
This distinction is the central firewall.
If a grammar assigns a role to 87.802% of tokens, that is coverage.
It does not mean 87.802% of translations are known to be correct.
Correctness requires an external truth signal.
A bilingual text.
A secure crib.
A known source.
A demonstrable physical label.
Or some other anchor that distinguishes one semantic reading from rivals.
The preprint itself recognises this distinction when it says the 98.3914% benchmark figure is not whole-manuscript lexical accuracy.
High structural coverage can prove that a grammar is productive and coherent. It cannot by itself prove that the English words attached to those roles are historically correct.
Why Benchmark Membership Matters
The 1,609-line translation benchmark is described as fully constrained.
That means the lines were selected because the frozen grammar could constrain them strongly enough for the project’s translation standard.
This is defensible if stated clearly.
It also means benchmark performance cannot be interpreted as a random sample of the entire manuscript.
Harder unresolved lines are partly outside the benchmark by design.
The scientific question becomes:
how was “fully constrained” defined before the English rendering, and would an independent analyst select the same lines under the same rules?
The Frozen Grammar Is the Most Important Claim
Voynich translation systems often fail because rules change whenever a passage resists.
One glyph means X here.
Something else there.
One ending is grammatical until it blocks a desired sentence.
Then it becomes optional.
The Dominik project explicitly tries to avoid this by freezing its compositional grammar and role lexicon before adversarial replication.
That is exactly the right methodological instinct.
The value of the claim depends on how complete that freeze really is.
- Were role definitions fixed?
- Were semantic gloss families fixed?
- Were recursion rules fixed?
- Were exception classes fixed?
- Were benchmark-selection rules fixed?
- Were English rendering conventions fixed?
The more of these were frozen prospectively, the stronger the evidence.
Adversarial Replication on RF1b Matters
One of the strongest aspects of the deposit is the attempt to move away from one canonical transcription.
The project reports post-freeze replication using the RF1b reference transliteration and alternate uncertain-space treatments.
The principal role architecture reportedly survives.
Exact whole-corpus coverage falls from 87.802% under ZL3b to 86.065% under RF1b.
This is scientifically healthy.
A real structure should survive reasonable representation changes.
A fragile solution often collapses when spaces or glyph readings move slightly.
But survival of role architecture still does not independently validate the English semantics attached to each role.
It validates robustness of the structural parser more directly.
The 4,258/4,258 Replication Result
The preprint reports that one replication run reproduces 4,258 of 4,258 canonical recursive occurrences while preserving frozen negatives.
This is an important reproducibility claim.
If independently replicated, it shows the parser is not relying on hidden manual intervention for those recursive structures.
Again, the evidence level must be named precisely.
It supports:
- deterministic structural recurrence;
- rule stability;
- reproducibility of the parser.
It does not automatically support:
- historical lexical identity;
- English word choice;
- source-language identification.
What Is a Functional-Role Interpretation?
This phrase is essential to understand.
A token may be assigned a structural role such as:
- action-like;
- material-like;
- location-like;
- quantity-like;
- connector-like;
- modifier-like;
without knowing the exact dictionary word.
This is familiar in linguistic analysis.
We can identify that an unknown word behaves like a verb before knowing whether it means “mix”, “heat”, “wash” or “carry”.
The translation claim becomes stronger only when role assignment and lexical choice are kept distinct.
Opaque material in brackets is therefore a methodological strength if it prevents role certainty from becoming invented lexical certainty.
Why Readable English Can Be Misleading
English is forgiving.
If we know a sequence contains action + object + location + modifier, many fluent English sentences can be produced.
For example:
- place the herb in the vessel;
- mix the material in the container;
- set the preparation beside the liquid.
All may fit the same broad role skeleton.
Therefore fluency of the English output is weak evidence by itself.
The critical issue is lexical uniqueness.
Why this English noun rather than another noun in the same role?
Why this action rather than another action?
What external observation forces the choice?
The Source-Language Question Remains Open
The deposit explicitly does not claim a recovered historical language.
Northwest Caucasian / Abkhaz–Abaza is retained only as a structural analogue.
This is an important boundary.
A grammar can resemble the slot architecture of a language family without being that language.
Typological similarity is not genealogy.
It is also not geography.
The system therefore cannot currently use its English output to claim that Voynichese is Abkhaz, Abaza or another Northwest Caucasian language.
Historical Phonology Is Still Missing
A complete decipherment normally connects written units to a language system deeply enough to explain sound, morphology, syntax and lexicon.
The preprint explicitly does not claim historical phonology.
That means the visible glyphs are not yet mapped to a secure sound system.
This is not fatal to every kind of translation.
Some writing systems can be translated functionally before pronunciation is fully known.
But it limits how strongly the result can be called a recovered language decipherment.
The Phonotactics Problem remains open.
Where Is the External Crib?
A powerful grammar can discover internal regularity.
External anchors are what turn roles into semantics.
The manuscript currently lacks a universally accepted Rosetta Stone.
No plant has a secure original-language label.
No zodiac ring supplies a secure bilingual name list.
No f116v phrase has become an accepted gloss.
Without such anchors, semantics must be inferred from internal role behaviour and contextual interpretation.
That is exactly where multiple coherent grammars can potentially compete.
See The Crib Problem.
The Strongest Test: Competing Frozen Grammars
One grammar producing coherent English is evidence.
Two incompatible grammars producing equally coherent English would be a problem.
The gold-standard next experiment is therefore competitive.
- Freeze the Dominik grammar and lexicon.
- Build independent alternative grammars without using its semantic outputs.
- Evaluate all systems on the same held-out lines.
- Use pre-declared structural and external criteria.
- Ask which model predicts newly revealed evidence most accurately.
The correct translation should not merely be possible.
It should increasingly become difficult for rival interpretations to survive.
Blind Lexical Anchoring Would Be Transformative
Imagine the grammar predicts, before seeing an image, that one token family should denote a container class.
Then a held-out page contains that family only beside vessel drawings.
Repeat this across several independent semantic classes.
Now the semantics are beginning to acquire external support.
The key is temporal order.
Prediction before inspection is stronger than explanation after inspection.
This principle should govern every future translation claim.
The Visual Classifier Could Become a Useful Independent Channel
The new Visual Classification Problem creates an interesting future control.
A translation grammar should predict document roles that correlate with visual page structure.
If a procedural grammar identifies ingredient-like, action-like or container-like roles, those distributions should interact with image families in stable ways.
The visual model must not be trained on the translation labels.
The translation grammar must not be tuned to the visual outputs.
Then cross-channel convergence becomes genuinely independent evidence.
Why the 201 Opaque Positions Matter
A weak translation system tends to erase uncertainty.
Every token receives a meaning because the method cannot admit failure.
The controlled benchmark reportedly retains 201 positions as opaque.
That is a healthy design choice.
But we should ask how those opaque positions are distributed.
- Do they cluster in rare glyphs?
- In one Currier regime?
- At uncertain spaces?
- In labels?
- At line boundaries?
- In specific hands?
Failure geography can reveal exactly where the grammar is incomplete.
What Would Count as Independent Replication?
Running the author’s code and reproducing the same outputs is reproducibility.
That is valuable.
Independent replication goes further.
- A separate team rebuilds the parser from the published rules.
- They use independently prepared transcription data.
- They recover the same structural roles.
- They make the same held-out predictions.
- They test semantic anchors the original project did not select.
- They report where their implementation disagrees.
That is the threshold at which a private coherent system begins to become a public scholarly result.
What the 1,609-Line Translation Claim Does Not Prove
- It does not prove the Voynich Manuscript has been completely deciphered.
- It does not prove 98.3914% lexical accuracy.
- It does not recover historical phonology.
- It does not provide a complete root dictionary.
- It does not identify Northwest Caucasian as the source language.
- It does not supply a secure external crib.
- It does not establish that every fluent English rendering is uniquely forced.
- It does not yet represent peer-reviewed scholarly consensus.
What It Can Give Us
- A large-scale explicit translation claim rather than isolated word readings.
- A frozen grammar that can be tested prospectively.
- Published coverage and opacity accounting.
- Alternate-transcription stress testing.
- Reproducibility files, manifests and checksums.
- A clear distinction between role architecture and historical phonology.
- A high-value target for adversarial independent replication.
Reader Checklist
- Is the claim structural coverage or lexical accuracy?
- How were the 1,609 lines selected?
- Were selection rules frozen before translation?
- Which grammatical rules were frozen?
- Which semantic glosses were frozen?
- Does the system survive alternate transcription?
- What happens to uncertain spaces?
- Where do opaque positions cluster?
- What external evidence forces the lexical choices?
- Could an incompatible grammar produce comparably coherent English?
- Has an independent team rebuilt the method?
- Does the translation predict new evidence prospectively?
Research Foundations
- Matthew Dominik — The Voynich Manuscript: A 1,609-Line English Translation with a Cross-Validated Procedural Grammar (August 2026 preprint).
- ResearchGate mirror and preprint record.
- What a Real Voynich Decipherment Must Survive.
- The Crib Problem.
- The Joint Benchmark Problem.
The Final Idea
A 1,609-line translation claim should not be dismissed because it is ambitious.
It should not be accepted because it is large.
Its value lies in how much of the method has been frozen, exposed and made vulnerable to replication.
The next phase is not rhetorical.
It is adversarial.
If the controlled translation is real, independent researchers should increasingly fail to find alternative grammars that explain the same held-out structure equally well.