Voynich theories are unusually good at explaining pictures after the pictures are already visible.
A strange plant appears.
A token on the page is declared to mean “leaf”.
Another is declared to mean “root”.
The illustration now seems to confirm the translation.
This is one of the oldest traps in the manuscript.
Once the answer is visible, the text can be interpreted toward it.
The cure is brutally simple.
Make the prediction before looking at the picture.
A 2026 computational project led under the Workwrite-Niidome repository did exactly this in two pre-registered blind-prediction rounds. Twenty previously unanalysed folios were selected in two groups of ten. Text-derived predictions were recorded first. Only then were the Yale facsimile illustrations examined and scored.
The result is valuable not because everything worked.
It did not.
The same project later expanded five promising morpheme–feature associations from a 30-page training set to all 112 herbal pages. Four of the five failed to replicate at full scale.
That failure is one of the most useful results in the entire programme.
Quick Read
- The 2026 Workwrite-Niidome project reports two pre-registered blind prediction tests on previously unanalysed herbal folios.
- Predictions were recorded before illustrations were examined.
- The project reports that several initial morpheme–illustration claims weakened or failed under blind testing.
- One early claim associating sh with roots was explicitly refuted after poor blind-test performance.
- The project later retested five top morpheme–feature associations from a 30-page training set across all 112 herbal pages.
- Four of those five associations failed to replicate at full scale.
- Only one association in that set—reported for ty and thin/linear leaves—survived the project’s full-scale test under its chosen criteria.
- Blind prediction is stronger than retrospective fit because the target evidence cannot influence the prediction.
- Full-scale replication is stronger again because small-sample patterns often disappear when exposed to the whole section.
- No blind-prediction success by itself establishes a translation; it establishes predictive association under the tested representation.
- Failure reporting should be treated as evidence quality, not embarrassment.
Why Voynich Is So Vulnerable to Retrospective Fitting
The manuscript contains thousands of words and hundreds of illustrations.
That creates an enormous search space.
If a researcher looks long enough, some word family will be frequent on pages with broad leaves.
Some other family will appear on pages with large roots.
Another will cluster near stars.
The danger grows when:
- many candidate morphemes are tested;
- many visual features are available;
- definitions can change after inspection;
- partial matches are scored generously;
- negative results disappear from publication.
Blind prediction removes one of those freedoms.
The text interpretation has to commit before the image is allowed to answer.
What Pre-Registration Adds
Blindness alone is helpful.
Pre-registration is stronger.
Before examining the held-out target, record:
- which token or morpheme is being tested;
- what visual feature is predicted;
- how success will be scored;
- what counts as partial;
- what chance baseline will be used;
- what result would falsify the claim.
This creates a receipt.
After the image is revealed, the rules cannot quietly move to accommodate it.
Blind Test 1: Why Failure Was More Valuable Than the Hits
The project reports an overall first-round score of 67.6% when partial predictions are counted as half credit, against an estimated chance range around 60–65%.
That margin is not overwhelming.
More importantly, one specific claim performed badly: the proposed sh → root association.
The project reports positive predictions at 25% and negative predictions below a useful threshold, prompting the claim to be explicitly refuted.
This is exactly what a blind test should do.
It should be capable of killing a favourite interpretation.
If a test cannot force a theory to lose something, it is not much of a test.
The Full-Scale Replication Was Harder
Small samples are dangerous.
The original project used a 30-page training set to identify promising associations between morphemes and visual features.
Then the top five were retested over all 112 herbal pages.
Four failed to replicate.
That is a major methodological result.
It shows how easily a plausible visual-semantic pattern can emerge in a small exploratory subset and disappear at scale.
It also tells us that Voynich illustration semantics are especially vulnerable to multiple-comparison effects.
Why 4 of 5 Failed Associations Should Be Published Prominently
There is a temptation to headline the one survivor.
That would be the wrong lesson.
The stronger lesson is the attrition rate.
Exploratory correlation generated five exciting candidates.
Full-scale replication eliminated four.
That means the discovery pipeline is functioning.
It also means any unreplicated Voynich image association deserves a heavy discount until it survives comparable exposure.
The Surviving ty Association Is Still Not a Translation
The project reports one association—ty with thin/linear leaves—surviving its full-scale test with Fisher p=0.008 and a reported odds ratio of 18.
This is interesting.
It does not establish that ty literally means “thin leaf”.
Possible alternatives remain:
- the token family marks a plant class correlated with thin leaves;
- the association reflects one scribal or Currier regime;
- the feature taxonomy captures another property correlated with leaf shape;
- the result is a surviving false positive that requires independent replication.
Predictive association is stronger than resemblance.
It is still not lexical identity.
Chance Baselines Are Harder Than They Look
If 70% of Voynich herbal pages have prominent leaves, predicting “prominent leaves” on every page achieves 70% accuracy without reading one glyph.
This is why raw hit rate can be misleading.
A good blind test needs feature base rates.
It should also report:
- sensitivity;
- specificity;
- positive predictive value;
- negative predictive value;
- odds ratios or effect sizes;
- permutation or held-out null distributions.
The project’s later full-scale analysis moves in this direction, which is substantially stronger than eyeballing successful pages.
The Multiple-Testing Problem
Voynich invites thousands of tests.
Twenty words × eight visual features already produces 160 comparisons.
At a conventional p<0.05 threshold, several apparently significant results can appear by chance.
The project reports a later 168-combination word–illustration analysis in which only one survived Bonferroni correction—and in the wrong direction for the hypothesised meaning.
That result should make every casual “this word appears beside roots” claim much more cautious.
Prospective Prediction Is Better Than Retrospective Explanation
The hierarchy of evidence is simple.
- Retrospective resemblance: after seeing the image, a token is said to fit it.
- Retrospective correlation: token frequency is compared with an already defined image feature.
- Held-out replication: the relationship is discovered in one set and tested in another.
- Prospective blind prediction: the text predicts an unseen image before inspection.
- Independent prospective replication: another team repeats the prediction under frozen rules.
Voynich claims should be ranked accordingly.
Why Image Identification Itself Can Leak
A researcher may believe a folio is unseen while still knowing its transcription, section or famous page number.
Those clues can leak visual expectations.
A stronger blind protocol therefore separates roles.
- One person selects target folios.
- Another receives text only.
- Predictions are timestamped.
- A third person reveals and scores images.
Automated random target selection can strengthen this further.
The Scoring Rubric Must Be Frozen Too
“Partial” is a dangerous word.
If a prediction says “large underground structure” and the page shows a large bulb, is that a hit?
What if it shows a broad root?
A tuber?
A pot?
The more flexible the scoring after reveal, the more blindness is weakened.
Therefore each prediction should have a rubric written before the target is opened.
A Real Translation Should Predict More Than Image Features
Blind image prediction is powerful because Voynich is illustrated.
It should not become the only external test.
A translation mechanism should also predict:
- labels on held-out diagrams;
- document-role changes;
- quantitative structure where independent counts exist;
- known zodiac/month relations;
- physical container or plant-fragment associations;
- cross-page lexical reuse.
Different external channels make collusion between theory and target less likely.
The Visual Classification Work Creates Another Blind Channel
The new Visual Classification Problem shows that image features can be represented computationally without reading Voynichese.
This creates an attractive future protocol.
- A text theory predicts a visual property of held-out pages.
- An independent image model scores those pages without seeing the text theory.
- The two outputs are compared automatically.
This could reduce human scoring flexibility significantly.
Blind Prediction and the Crib Problem
A secure crib is the ultimate external prediction test.
A theory assigns a meaning before the anchor is revealed.
The historical external source then confirms or contradicts it.
Voynich currently lacks a universally accepted crib, which makes prospective visual and structural testing especially valuable in the meantime.
See The Crib Problem.
Failure Geography Is Evidence
When a prediction system fails, ask where.
- One Currier regime?
- One proposed hand?
- One plant style?
- Long labels?
- Rare token families?
- Pages with damaged imagery?
- Line-final forms?
A failure map can reveal hidden variables even when the semantic claim collapses.
This is one reason failed Voynich predictions should be preserved rather than deleted from the research record.
What Blind Prediction Does Not Prove
- One successful prediction does not prove translation.
- A high hit rate does not matter without a fair chance baseline.
- Blind image association does not establish exact lexical meaning.
- Small blind samples do not substitute for full-scale replication.
- A surviving association does not automatically generalise beyond herbal pages.
- Failure of one morpheme meaning does not refute every structural property of that morpheme.
What It Can Give Us
- A defence against retrospective storytelling.
- A way to force semantic hypotheses to make risky predictions.
- A clean distinction between exploratory discovery and confirmatory testing.
- A mechanism for publishing failure as useful evidence.
- A bridge between text-only hypotheses and independent image evidence.
Primary School: Guess Before Turning the Card
Put a picture face down.
Give the child a coded clue.
Ask for a written prediction before the card is turned.
The child immediately learns why guessing after seeing the answer is not prediction.
Secondary School: Discovery Set and Test Set
Use twenty pages to discover a candidate relationship.
Freeze it.
Then test on twenty unseen pages.
Students see how many exciting small-sample relationships disappear under replication.
JC and Adult Readers: Build a Prospective Protocol
- Declare the theory.
- Choose unseen targets by random rule.
- Hide target images.
- Record text-derived predictions with timestamps.
- Freeze scoring criteria.
- Reveal images.
- Score automatically or independently.
- Compare against feature base rates.
- Replicate on the full relevant section.
Reader Checklist
- Was the target genuinely unseen?
- Was the prediction recorded before reveal?
- Was the scoring rubric frozen?
- What is the chance/base-rate benchmark?
- How many predictions were attempted?
- Were failed predictions reported?
- Was multiple testing corrected?
- Was the claim replicated on the full section?
- Did the interpretation change after failure?
- Has another team repeated the test?
Research Foundations
- Workwrite-Niidome — Voynich Manuscript Multi-Agent Structural Analysis, Final Paper v3 (2026 research release).
- Public repository and analysis record.
- The Control Problem.
- What a Real Voynich Decipherment Must Survive.
The Final Idea
Voynich research does not need fewer bold ideas.
It needs bolder ways to let those ideas fail.
The cleanest prediction is the one written down before the manuscript is allowed to show us the answer.