VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Joint Benchmark Problem: Why Passing Seven Voynich Tests Can Still Mean Failing the Manuscript

Voynich theories have always been good at winning one game.

Zipf.

Entropy.

Word length.

Local repetition.

Currier A/B.

Position effects.

A mechanism reproduces one of them and suddenly sounds much closer to the manuscript.

The problem is that the Voynich Manuscript is not one statistic.

It is a conjunction.

In May 2026, Yang Ou proposed a “Statistical Turing Test” for Voynich: eight quantitative properties applied simultaneously to eighteen model variants spanning nine generator families. No tested model passed all eight. The strongest section-stratified interpolated bigram model passed seven.

That result is a preprint and should be treated as a methodological experiment rather than settled manuscript science.

But the core idea is extremely important.

A Voynich model should not be judged by the statistic it was built to imitate. It should be judged by the independent properties it still reproduces when all of them are asked at once.

Quick Read

  • The May 2026 Yang Ou preprint defines an eight-metric Statistical Turing Test for Voynich-like generated text.
  • Eighteen model variants across nine model families were evaluated over repeated runs.
  • The metrics include Zipf slope, character conditional entropy, word conditional entropy, standardised type-token ratio, word-length distribution, bigram NPMI, positional vocabulary divergence and Currier-dialect replicability.
  • No tested model passed all eight criteria simultaneously.
  • The strongest reported model, a section-stratified interpolated bigram sampler, passed seven of eight.
  • The Cardan-grille simulation passed only one of eight in that benchmark.
  • GPT-2 fine-tuned on the manuscript also passed only one of eight under the paper’s criteria.
  • The hardest property in the benchmark was Zipf slope: all tested models produced a steeper rank-frequency slope than the manuscript.
  • The metric most resistant to stationary one-vocabulary generators was positional vocabulary divergence.
  • The benchmark therefore argues that section structure matters, but it does not prove that the winning seven-of-eight model is historically correct.
  • The deeper lesson is joint evaluation: models must survive properties they were not specifically tuned to reproduce.

Why One Successful Statistic Is Weak Evidence

Many different mechanisms can reproduce Zipf-like frequency.

Many can reproduce low entropy.

Many can reproduce a narrow word-length distribution.

That means each statistic by itself has low discriminatory power.

The question is not:

Can model X look Voynich-like under metric M?

The stronger question is:

Can one fixed model occupy the Voynich region across many independent dimensions at the same time?

This is the difference between resemblance and constraint satisfaction.

What the Eight Metrics Try to Measure

Ou’s benchmark combines several familiar Voynich properties rather than trusting any one of them.

  • Zipf slope: how quickly token frequency falls with rank.
  • Character conditional entropy: how predictable characters are from local context.
  • Word conditional entropy: how predictable token succession is.
  • Standardised type-token ratio: vocabulary diversity after length normalisation.
  • Word-length distribution: the shape of token lengths.
  • Bigram NPMI: strength of recurring adjacent character associations.
  • Positional vocabulary divergence: how vocabulary changes across manuscript position or section.
  • Currier replicability: whether generated output can reproduce the A/B-like distinction.

These are not the only possible metrics.

The recent Edge, Hapax, Unit-Scale and Directional Dissociation results already suggest additional tests that future batteries should include.

That is a strength of the benchmark idea.

The test can grow as the manuscript’s known constraints grow.

Why Seven Out of Eight Is Still a Failure

Seven of eight sounds excellent.

For historical mechanism identification, it is not enough.

If the eighth property is independent and genuine, failing it means the model is missing part of the manuscript’s architecture.

A model can be useful while still being wrong.

It can reveal which features are easy to reproduce.

It can identify the hard residual.

It can become a control.

But seven of eight should not be rhetorically converted into:

therefore this is how Voynich was made.

The missing eighth property is exactly where the historical mechanism may be hiding.

Why Section Stratification Helps So Much

One of the preprint’s strongest findings is methodological rather than mechanistic.

A model drawing from one stationary vocabulary distribution performs much worse than a model that conditions generation on manuscript section.

The best L8b model is section-stratified.

That echoes what Voynich researchers already know qualitatively.

Herbal pages do not distribute vocabulary exactly like Quire 13.

Currier A and B are not interchangeable.

Labels and prose differ.

The manuscript is not one homogeneous bag of tokens.

The benchmark translates that into a model-design lesson:

global fit can fail because the manuscript is locally structured.

Why That Still Does Not Prove Sections Are the True Hidden Variable

Section stratification is useful.

It may also be proxying for other things.

  • scribe;
  • Currier regime;
  • source exemplar;
  • chronology;
  • document role;
  • image type;
  • production batch.

If several of these covary with conventional section labels, a section-aware generator can improve without discovering the actual historical cause.

This is the same hidden-variable problem explored in The Drift Problem.

The Cardan-Grille Result: A Quantified Weakness, Not Universal Refutation

The benchmark reports that its Cardan-grille simulation passes only one of eight metrics.

This is a useful quantitative result against that tested implementation.

It should not be over-expanded into:

no grille-like or tabular generation mechanism could ever produce Voynich.

Implementation matters.

Parameterisation matters.

Historical variants matter.

The correct claim is bounded:

the tested Cardan-grille simulation does not reproduce the joint benchmark well.

The Zipf Problem Is Surprisingly Hard

Zipf’s law is often treated as one of the easiest Voynich statistics to reproduce.

Ou’s benchmark complicates that story.

The paper reports that every tested model produces a rank-frequency slope steeper than the manuscript’s baseline.

Some models fall within the paper’s tolerance window, but the systematic direction remains.

The manuscript therefore appears unusually resistant to concentration in the most frequent token types relative to the tested generators.

The preprint interprets this as evidence for an “anti-concentration” mechanism beyond simple frequency-weighted sampling.

That interpretation is interesting.

The safer result is:

the tested generators struggle to reproduce the manuscript’s particular balance between common and rare types.

This Connects Directly to the Hapax Problem

A flatter Zipf slope and a hapax-rich vocabulary are related manifestations of the same broader tension.

Voynich has common forms.

But it also keeps producing rare forms.

A generator that over-concentrates on popular tokens steepens the frequency curve and suppresses singleton growth.

A generator that produces too much novelty loses Voynich’s internal regularity.

Joint benchmarking forces the model to balance both.

See The Hapax Problem.

A Benchmark Can Overfit Too

There is an irony here.

A benchmark designed to prevent model overfitting can itself become overfitted to the current research fashion.

If we choose eight metrics because they are famous, a model can eventually be tuned to those eight.

Then it passes the battery while failing a ninth property no one included.

This is why the recent 2026 results matter so much.

Edge coupling.

Directional dissociation.

Open vocabulary.

Unit scale.

Each new independent property can become an out-of-sample test for models developed before it was measured.

The best Voynich benchmark is not a frozen exam. It is a growing wall of independent constraints.

Training Metrics and Testing Metrics Must Be Separated

This principle should be explicit.

If a generator is tuned until its Zipf slope matches Voynich, Zipf is no longer a clean validation metric.

If it is tuned on word length, word length is no longer independent validation.

A stronger workflow is:

  1. choose a subset of properties for fitting;
  2. freeze the model;
  3. evaluate on different properties;
  4. evaluate on held-out pages or quires;
  5. publish failures as well as successes.

This is exactly the logic behind the user’s broader evidence discipline: one clue should not be allowed to do the work of five.

A Historical Mechanism Has an Extra Benchmark

A statistical generator can pass all eight tests and still fail history.

Imagine a modern neural model that generates perfect Voynich-like text.

That proves statistical capability.

It does not prove a fifteenth-century scribe had access to that mechanism.

A historical theory needs a second class of tests:

  • manual executability;
  • materials and tools;
  • scribal fluency;
  • page planning;
  • correction behaviour;
  • plausible exemplars or workflows.

Statistical indistinguishability is therefore necessary for some mechanism claims.

It is never sufficient for historical attribution.

The Benchmark Should Respect Representation Uncertainty

Every metric in the battery is computed on a representation of the manuscript.

Change glyph segmentation and character entropy changes.

Change uncertain spaces and word length changes.

Merge multi-symbol units and bigram statistics change.

A robust benchmark should therefore report sensitivity to:

  • multiple EVA transcriptions;
  • uncertain-space policies;
  • compound-glyph policies;
  • different section partitions.

Ou’s preprint explicitly compares several EVA transcriptions and reports that metric variation remains within its tolerance bands for the tested versions, which is a useful robustness check.

What Passing the Benchmark Would Actually Mean

If a future model passes all eight metrics, what has been proven?

Only this:

under the chosen representation and tolerances, the generated output is statistically similar to Voynich across those eight measured properties.

It would not prove:

  • same semantics;
  • same historical mechanism;
  • same production process;
  • same underlying language;
  • same page-level multimodal organisation.

A Turing-style benchmark tests behavioural imitation.

Historical explanation requires causal evidence.

What the Joint Benchmark Problem Does Not Prove

  • It does not prove the eight chosen metrics are complete.
  • It does not prove the seven-of-eight model is close to the historical mechanism.
  • It does not eliminate all Cardan-grille variants.
  • It does not eliminate all neural models.
  • It does not prove section labels are the true causal variable.
  • It does not convert preprint tolerances into universal Voynich thresholds.

What It Can Give Us

  • A standard for simultaneous rather than cherry-picked evaluation.
  • A way to compare very different generator families under one battery.
  • A way to identify which Voynich properties remain hardest to imitate.
  • A discipline for separating fit metrics from held-out metrics.
  • A foundation for adding new 2026 constraints as out-of-sample tests.

Reader Checklist

  1. How many metrics were tested?
  2. Which were used to tune the model?
  3. Which were held out?
  4. Were all metrics passed simultaneously?
  5. Were tolerance bands declared before evaluation?
  6. Does the result survive alternate transcriptions?
  7. Were section effects modelled?
  8. Were new post-hoc metrics added only after seeing failures?
  9. Does the model also survive Edge, Hapax, Unit-Scale and Directional tests?
  10. Is statistical imitation being confused with historical mechanism identification?

Research Foundations

The Final Idea

The Voynich problem is full of theories that are right about something.

The hard part is finding one that is not wrong about the rest.

A model that passes seven tests has taught us something. A model that fails the eighth has also told us where to look next.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading