VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Prefix–Core–Suffix Generation Problem: How Much of Voynich Vocabulary Can a Three-Part Generator Reconstruct?

A descriptive grammar becomes much more interesting when it can generate.

Voynich tokens have long been described as if they contain preferred beginnings, middles and endings.

But description is cheap.

A serious model should be able to construct legal-looking tokens from those parts and recover the actual vocabulary better than chance.

That is the job of Youngsan Chang’s 2026 prefix–core–suffix generator.

The integrated reproduction package reports two headline numbers for the PCS model:

  • 97.02% matching of observed token types under the model’s coverage definition;
  • 85.25% token-mass coverage.

The same package reports random baselines of 5.07% and 18.50% respectively.

Those are large differences.

The result is still preprint-level and tightly linked to the representation used to define prefix, core and suffix.

The question therefore is not whether the generator is impressive.

It is.

The question is whether the generator has discovered a genuine Voynich production grammar or simply become very good at recombining units defined from the same vocabulary it is asked to explain.

Quick Read

  • Chang’s April 2026 integrated package includes an explicit prefix–core–suffix token-generation model.
  • The package reports 97.02% matching of observed Voynich types and 85.25% token-mass coverage.
  • Random baselines are reported at 5.07% type matching and 18.50% token-mass coverage.
  • The result strongly supports the existence of reusable token components under the study’s decomposition.
  • It does not prove those components are linguistic morphemes.
  • It does not prove the PCS model is the historical mechanism.
  • A high coverage score can be inflated if the component inventory is learned from the same corpus being reconstructed.
  • The strongest next test is held-out generation: learn the component inventory on one set of quires, freeze it, and predict unseen token types in other quires.
  • A historical mechanism must also reproduce frequency, line position, Currier variation, hapax growth, edge coupling and local sequence behaviour.

What Does “97.02% Matching” Mean?

The number sounds almost like translation accuracy.

It is not.

The model generates token forms from reusable prefix, core and suffix inventories.

The matching percentage measures how much of the observed type vocabulary falls inside the model’s legal generative space under the paper’s rules.

That is a structural coverage result.

It says:

the visible vocabulary can be reconstructed extensively from a compact component system.

It does not say:

97.02% of Voynich words have been translated correctly.

Token-Mass Coverage Is a Different Statistic

Type coverage treats every distinct token form equally.

Token-mass coverage weights frequent forms more heavily because each occurrence counts.

A model can cover nearly all frequent tokens while missing many rare types.

Or it can cover many rare forms while missing important common ones.

The reported 85.25% token-mass coverage therefore complements the 97.02% type figure.

It suggests the model’s legal space is not merely populated by obscure edge cases.

Still, the gap between the two metrics is informative.

Any omitted 14.75% of token mass deserves to be mapped rather than treated as noise.

Why the Random Baseline Matters

A flexible component generator can cover many forms simply because recombination creates a large legal vocabulary.

That is why random baselines are essential.

The package reports only 5.07% type matching and 18.50% token-mass coverage under its random control.

The difference implies that the observed PCS structure is not reproduced by arbitrary recombination under the chosen control.

The result therefore supports constrained combinatorics.

It does not yet tell us what historical mechanism imposed the constraints.

The Generator Is Related to Word Families, but It Owns a Different Job

Word Families establishes the observation that Voynich tokens keep looking almost alike.

The PCS generator asks whether that resemblance can be formalised into productive rules.

Observation:

these words look related.

Generator:

these reusable components can reconstruct most of the observed vocabulary under explicit legal combinations.

That is a substantially stronger claim.

The Circularity Risk

Every generative vocabulary model faces one obvious danger.

If the components are learned from the complete observed vocabulary and then recombined to reconstruct that same vocabulary, part of the success can be circular.

This does not make the model useless.

It tells us which next experiment matters most.

Prospective vocabulary prediction

  1. Learn prefixes, cores, suffixes and legal dependencies on a training set.
  2. Freeze every inventory and rule.
  3. Hide several quires.
  4. Generate the legal unseen vocabulary.
  5. Ask how many held-out token types fall inside the predicted space.
  6. Measure how many predicted forms never occur.

The false-positive count is as important as the recovered count.

Overgeneration Is the Hidden Cost

A model can achieve excellent coverage by allowing far too many possible tokens.

Imagine a generator that produces every possible six-character string.

It will “cover” every real six-character Voynich token.

It explains nothing.

The important paired metric is therefore precision:

of all forms the PCS model declares legal, how many are actually attested or become attested in held-out data?

Coverage without overgeneration accounting can make a generator look stronger than it is.

The Core–Suffix Dependency Gives the Generator Teeth

The previous article established why one internal dependency matters.

If every core could take every suffix, the PCS model would generate a huge combinatorial space.

Core-conditioned suffix selection shrinks that space.

That makes the model more falsifiable.

See The Core–Suffix Dependency Problem.

Can the Generator Explain Hapaxes?

Voynich contains a large singleton tail.

A good token generator must therefore produce novelty.

But it must produce disciplined novelty.

If the PCS grammar is real, many hapaxes should be rare legal combinations of common subcomponents.

That creates a test.

  • Are singleton forms disproportionately composed of common components?
  • Do they obey the same core→suffix rules?
  • Do they occupy the same positional regimes?
  • Does the generator predict their structural families before seeing them?

The Hapax Problem is therefore a direct stress test for the PCS model.

Can It Explain Currier A/B?

If the same component grammar operates across the manuscript, Currier A/B may arise through changed component frequencies or legal combinations.

If separate inventories are required, the model begins to look more like several related grammars.

Both are possible.

The important point is to predeclare which elements are shared and which vary.

Otherwise every Currier mismatch can be rescued by giving one regime its own local parameter set.

Can It Explain Line Position?

Voynich legal tokens do not distribute uniformly across lines.

Some forms concentrate at starts.

Others prefer endings.

A token generator that ignores line position may reconstruct vocabulary while still failing the manuscript’s page grammar.

The PCS model therefore becomes more historically interesting if positional state predicts which prefixes, cores or suffixes are legal.

This links directly to The Line as a Unit.

Can It Explain Labels?

Labels and running prose use overlapping but non-identical vocabularies.

If one PCS grammar generates both, the model should explain the document-role difference through component probabilities or constraints.

If labels require another grammar, that distinction must be explicit.

Again, success at whole-corpus vocabulary reconstruction is not enough.

The generator must put the right forms in the right jobs.

Historical Interpretation: Morphology, Cipher or Notation?

The same PCS architecture can support several historical models.

  • Morphology: prefix and suffix are grammatical elements around a lexical root.
  • Cipher: the components are code groups with constrained compatibility.
  • Shorthand: components are reusable abbreviational chunks.
  • Notation: prefix/core/suffix encode metadata fields in a structured record.
  • Pseudo-text generation: components are template slots used to manufacture legal tokens.

The generator tells us that a compositional architecture can cover the vocabulary.

It does not tell us which historical meaning the slots had.

The Frequency Problem

A generator must reproduce more than membership in the legal vocabulary.

It must produce the right frequencies.

If a legal token occurs once in Voynich and 10,000 times in the synthetic corpus, the model has failed distributionally.

Therefore the next layer is probabilistic generation:

  • component frequency;
  • conditional component choice;
  • section state;
  • line state;
  • local sequence state.

This is where the PCS generator meets the Joint Benchmark Problem.

Local Sequence Is the Next Exam

Chang’s integrated package also reports that real local token sequences have lower bigram entropy and higher next-token prediction than line-internal shuffle baselines.

This means a vocabulary generator that chooses legal tokens independently is incomplete.

The manuscript appears to care which legal token follows which.

A historical generator therefore needs two layers:

  • construct legal tokens;
  • arrange them under local sequence constraints.

That is a much harder model.

The Best Possible Result Would Be Prospective

Imagine a new damaged folio fragment were discovered tomorrow.

The PCS grammar is already frozen.

Could it predict which new token types are legal before the fragment is transcribed?

That is the dream test.

We may never receive a new folio.

Held-out quires can simulate the same logic.

What the PCS Generator Does Not Prove

  • It does not prove Voynich words contain linguistic prefixes and suffixes.
  • It does not prove 97.02% decipherment.
  • It does not prove the historical writer used this generator.
  • It does not establish semantics.
  • It does not establish spoken language.
  • It does not prove all observed forms are generated from exactly three historical components.
  • It does not solve overgeneration unless legal-space precision is reported.

What It Can Give Us

  • An explicit compositional reconstruction of most observed token types.
  • A dramatic improvement over a declared random baseline.
  • A testable vocabulary-generation grammar.
  • A bridge between word-family observations and full sequence models.
  • A platform for prospective held-out prediction.

Reader Checklist

  1. How were prefix, core and suffix inventories learned?
  2. Was the same corpus used for discovery and evaluation?
  3. How large is the generated legal vocabulary?
  4. What is the false-positive rate?
  5. What is held-out type coverage?
  6. What is held-out token-mass coverage?
  7. Does the model reproduce hapax growth?
  8. Does it reproduce Currier variation?
  9. Does it reproduce line-position constraints?
  10. Does it reproduce local sequence entropy?
  11. Can parameters remain fixed across new quires?
  12. Is structural coverage being confused with historical meaning?

Research Foundations

The Final Idea

Voynich vocabulary may be enormous without being arbitrary.

A small set of reusable components can generate a surprisingly large surface world.

The PCS model makes that intuition executable.

The next proof is not that the generator can rebuild what it has seen. It is that, once frozen, it can predict what it has not.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading