Voynich has a habit of looking ordinary from far away and strange up close.
Frequency laws look familiar.
Some entropy curves look familiar.
Vocabulary growth can look familiar.
Then we zoom into the token.
Beginnings cluster.
Endings concentrate.
Internal units behave as though only some combinations are legal.
And the familiar language-like silhouette begins to hide an unfamiliar machine.
Youngsan Chang’s April 2026 multi-level statistical programme makes this contrast explicit. Its comparative predictability study reports that Voynich entropy and ambiguity curves are broadly natural-language-like, while suffix concentration and prefix-family organisation remain structurally divergent.
The result is preprint-level and should be replicated independently.
But the question it raises is fundamental.
How can the manuscript occupy a language-like statistical region globally while organising its visible tokens in ways that remain unlike ordinary alphabetic prose?
Quick Read
- Chang’s 2026 integrated package compares Voynich predictability with natural-language controls at several structural levels.
- The package reports broadly natural-language-like entropy and ambiguity curves.
- At the same time, suffix concentration and prefix-family organisation remain unusually structured.
- This combination is more informative than either result alone.
- Global language-like behaviour does not prove the visible token is an ordinary word.
- Local structural divergence does not prove the manuscript is non-linguistic.
- A transformed language, shorthand system, verbose cipher, notation or constrained generator can preserve some global language statistics while changing local architecture.
- The strongest future test is representation-sensitive: repeat the curves after changing glyph segmentation, uncertain spaces and learned multi-symbol units.
- A historical mechanism should explain both the familiar global curves and the unusual local constraints with one coherent process.
What Is an Ambiguity Curve?
The basic idea is simple.
At one level, a symbol or token can have many possible continuations.
As more context is revealed, some possibilities disappear.
A curve can track how uncertainty falls as the system receives more information.
Natural languages have characteristic predictability profiles because spelling, morphology, syntax and vocabulary are all constrained.
Voynich can be compared with those profiles without assigning meaning to any glyph.
That is useful.
It is also easy to overinterpret.
Different mechanisms can generate similar curves.
A Global Curve Is a Projection
Imagine compressing a city into one statistic.
Average building height might make Singapore and another city look similar.
Street geometry, zoning, transport, density and neighbourhood structure can still differ dramatically.
An entropy curve works the same way.
It summarises a system.
It does not reveal every mechanism that produced the summary.
This is why “Voynich has a language-like curve” should never be translated into “Voynich is ordinary prose in disguise.”
The Local Divergence Is the Interesting Part
Chang’s integrated summary highlights two local structures that remain unusual despite the language-like global curves:
- strong suffix concentration;
- organised prefix families.
That means the token is not simply a sequence whose local possibilities resemble ordinary alphabetic spelling.
Its edges appear to be organised into relatively constrained families.
This connects directly to the newer Core–Suffix Dependency Problem and Prefix–Core–Suffix Generation Problem.
Could Ordinary Language Still Do This?
Yes.
Natural language is not one thing.
Inflection-heavy languages can have strong ending classes.
Abjads can suppress vowels.
Abbreviation traditions can compress recurrent material.
Technical registers can restrict morphology.
A transformed language can therefore preserve some large-scale statistical behaviour while changing the visible token shape substantially.
The correct comparison is not only against modern English spelling.
It is against historically plausible transformed writing systems too.
Could a Cipher Preserve the Global Curve?
Also yes.
A cipher derived from natural-language plaintext can inherit some source statistics.
A verbose cipher can then reshape the local symbol architecture.
One plaintext unit may expand into several visible symbols.
Code groups may have legal beginnings and endings.
The result can look language-like globally while appearing unusually templated locally.
This is one reason the Workshop Cipher Problem remains relevant as a control.
Could Shorthand Produce the Same Split?
Yes.
Shorthand can preserve the statistical skeleton of language while compressing common sequences into recurrent signs or chunks.
That reduces local uncertainty and can create strong positional families.
The Shorthand Problem therefore supplies another historically plausible comparator.
Could a Generator Do It Without Meaning?
Yes again.
A constrained generator can be built to reproduce a language-like entropy profile.
Self-citation can create family structure and familiar frequency laws.
Slot-based generators can create concentrated suffixes.
This is why the ambiguity curve is not a meaning detector.
It is a constraint that every candidate generator must match jointly with other properties.
The Representation Can Change the Curve
Voynich statistics are calculated on a transliteration.
If one visible ligature is split into two EVA characters, local predictability increases artificially.
If uncertain spaces are treated as hard word boundaries, token-level ambiguity changes.
If recurring multi-symbol units are collapsed into one unit, character-level curves change again.
The stronger research programme therefore asks:
- Does the language-like global shape survive alternate transcriptions?
- Does it survive merged uncertain spaces?
- Does it survive fused bench characters?
- Does it survive learned unit representations?
- Does local divergence survive the same transformations?
Currier A/B Can Create a Mixture Curve
Pooling different regimes can produce a global curve no single regime actually has.
This is a standard mixture problem.
If Currier A and B have different local distributions, their combination can look smoother or more language-like than either separately.
Therefore ambiguity curves should be computed:
- whole manuscript;
- Currier A;
- Currier B;
- by proposed hand;
- by quire;
- by label/prose role.
The Drift Problem makes this especially important.
A Familiar Curve Can Be the Result of Several Hidden States
Suppose the manuscript contains three sub-systems.
Each has unusual local structure.
Mix them and the aggregate uncertainty curve can move toward a familiar language region.
This creates a warning for every whole-corpus statistic:
global normality can be produced by averaging local abnormality.
The Strongest Mechanism Must Explain Both Scales
A model that explains only the language-like curve is incomplete.
A model that explains only the prefix/suffix architecture is incomplete.
The historical system—whatever it was—produced both.
That is why the best mechanism tests are joint.
See The Joint Benchmark Problem.
What the Ambiguity Curve Problem Does Not Prove
- It does not prove Voynich is natural language.
- It does not prove Voynich is non-language.
- It does not prove ordinary words are the correct unit.
- It does not prove prefix and suffix regions are linguistic morphemes.
- It does not eliminate cipher, shorthand, notation or generation.
- It does not make one preprint curve a final manuscript invariant.
What It Can Give Us
- A multi-scale constraint instead of a single statistic.
- A warning against treating global language-likeness as local mechanism identity.
- A way to test transformed language, cipher, shorthand and generator models under the same evidence.
- A reason to stratify curves by Currier, hand, quire and document role.
- A direct bridge from entropy to token architecture.
Reader Checklist
- Which units define the curve?
- Which transcription is used?
- Are uncertain spaces preserved?
- Are Currier A/B pooled?
- What natural-language controls are used?
- Are historical shorthand and cipher controls included?
- Which local token structures diverge?
- Does the same mechanism explain global and local behaviour?
- Does the result survive alternative representations?
- Is “language-like” being promoted into “language identified”?
Research Foundations
- Youngsan Chang — Multi-Level Statistical Constraints in the Voynich Manuscript (2026 reproduction package).
- Chang 2026 component Paper 2.
- The Unit-Scale Problem.
- The Joint Benchmark Problem.
The Final Idea
Voynich may look familiar when averaged and unfamiliar when opened.
That is not a contradiction.
It is a clue about scale.
The right explanation must account for why the manuscript can resemble language statistically without behaving like ordinary written language at every structural level.