VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Core–Suffix Dependency Problem: Does the Middle of a Voynich Token Predict Its Ending?

A Voynich token does not behave like a bag of independent glyphs.

The beginning matters.

The middle matters.

The ending matters.

And, increasingly, the interesting question is not whether those regions differ.

It is whether one region predicts another.

In April 2026, Youngsan Chang’s multi-level statistical programme reported a particularly sharp result: when Voynich tokens are decomposed operationally into prefix, core and suffix regions, suffix choice is strongly conditioned on the core. The integrated reproduction package reports H(suffix|core)=0.8370 and a 79.11% top-1 suffix prediction accuracy.

The result is preprint-level evidence and depends on the decomposition used.

But if it survives independent replication, it tells us something important before translation.

The ending of a Voynich token may be constrained by the structure immediately before it far more strongly than a flat alphabetic model would lead us to expect.

Quick Read

  • Chang’s 2026 integrated reproduction package synthesises five component studies of Voynich structure.
  • Paper 3 focuses on dependency between an operationally defined token core and suffix.
  • The package reports conditional suffix entropy H(suffix|core)=0.8370.
  • It also reports 79.11% top-1 accuracy when predicting suffix from the core.
  • This does not prove the suffix is a linguistic suffix in the historical-language sense.
  • A similar dependency can arise from morphology, abbreviation, verbose cipher groups, shorthand, notation, templated generation or transcription choices.
  • The result is strongest when compared with baselines and when repeated across Currier regimes, hands and held-out quires.
  • It should be tested under alternative token decompositions rather than assuming prefix/core/suffix is the true historical segmentation.
  • A mechanism that generates Voynich tokens should reproduce this dependency without adding post-hoc local exceptions.

What Does “Core” Mean Here?

The word core can sound more historical than it is.

In this research context, it is an operational region inside the visible token.

That is useful.

It is not the same as saying:

this is the lexical root of a medieval word.

Voynich token structure has long invited layered descriptions because some glyph families cluster toward beginnings, others toward middles, and others toward endings. Stolfi’s older crust–mantle–core terminology, slot models, prefix families and suffix classes all describe this broad asymmetry from different angles.

The modern question is whether those regions are merely positional or causally/statistically coupled.

Conditional Entropy in Plain English

Entropy measures uncertainty.

If many suffixes are possible and the core tells us almost nothing, suffix uncertainty remains high.

If seeing the core sharply narrows the possible suffixes, conditional entropy falls.

So H(suffix|core)=0.8370 is not a translation.

It is a statement about predictability.

The middle region reduces uncertainty about the ending.

The 79.11% top-1 prediction figure expresses the same idea in a more intuitive form: under the study’s model and decomposition, the single most probable suffix associated with a core is correct roughly four times out of five.

Those are striking numbers.

The important work begins after we resist the urge to name the mechanism too early.

Could This Be Ordinary Morphology?

Yes.

Languages constrain endings.

A stem can select a declension class.

A verb class can favour one inflectional ending.

Phonology can make some suffixes legal after some stems and illegal after others.

So strong core→suffix dependence is compatible with language.

But ordinary morphology is only one mechanism family.

The dependence would become more linguistically persuasive if suffix classes also predicted grammatical environments, neighbouring token classes or stable semantic roles independently of the decomposition used to discover them.

Could This Be Phonotactics?

A spoken-language explanation can also create strong ending constraints.

Some endings may simply be easier or legal after certain stem-final sounds.

The newer Voynich literature includes claims of cross-word suffix behaviour and pronounceable structural classes, but those claims remain contested and preliminary.

The existing Phonotactics Problem owns the sound-system question.

This article asks a lower-level question:

whatever these regions are, how tightly does one constrain the other?

Could This Be a Cipher Table?

Absolutely.

A verbose or stateful cipher can generate visible groups whose endings depend on the code group chosen in the middle.

One component may select a table row.

Another may encode a homophone class.

A final component may close or check the group.

Such a system can create morphology-like structure without ordinary morphology underneath.

This is why The Workshop Cipher Problem and this core–suffix article should remain connected but separate.

One asks whether a cipher mechanism can reproduce the manuscript broadly.

This one asks whether a particular internal dependency must be reproduced.

Could This Be Shorthand?

Shorthand and abbreviation systems often compress predictable material.

A core may imply a conventional ending.

A written suffix may represent a broader phrase class rather than a phonetic ending.

This can create low conditional entropy without requiring one glyph per sound.

The Shorthand Problem is therefore another necessary control.

Could This Be a Generator Slot?

Yes.

A token generator can use rules such as:

  • choose one core family;
  • choose from only the suffixes legal for that core;
  • optionally add a prefix;
  • apply spelling variation.

That mechanism would naturally produce high core→suffix predictability.

This is why the next owner in this batch—the Prefix–Core–Suffix Generation Problem—matters.

A dependency statistic becomes much more informative when embedded in an explicit generator whose output can be compared with real Voynich tokens.

The Baseline Problem

Seventy-nine percent sounds strong.

Strong compared with what?

If one suffix occurs 78% of the time globally, then 79.11% prediction is barely informative.

If suffixes are balanced globally and core identity raises prediction to 79.11%, the result is much stronger.

Every future discussion should therefore report:

  • global majority-class baseline;
  • frequency-weighted random baseline;
  • held-out predictive accuracy;
  • performance by Currier regime;
  • performance by hand and quire.

The value of a predictive percentage lives in its contrast with a fair null.

The Decomposition Problem

Core and suffix are not independent discoveries if the segmentation procedure was designed using recurring suffix patterns.

This is not automatically invalid.

It means confirmatory testing must move outside the data used to define the decomposition.

A rigorous workflow is:

  1. define prefix/core/suffix rules on training data;
  2. freeze the rules;
  3. apply them to held-out quires;
  4. measure conditional entropy and prediction without redefining regions.

If the same dependency survives, confidence rises sharply.

Currier A/B Is a Necessary Stress Test

Voynich internal structure varies across Currier regimes.

A core–suffix relationship that exists only in one regime is still real, but it is not universal.

A relationship that survives both with similar strength suggests a deeper common grammar.

A relationship that changes systematically may itself become a discriminator between regimes.

This is where the Currier A and B and Drift Problem owners become essential controls.

Hand and Quire Can Confound the Result

Suppose one proposed scribe prefers core family X and suffix Y.

Another prefers core family Z and suffix W.

Pool them and the corpus shows strong core→suffix dependence even if each scribe is simply using a different local convention.

That is still manuscript structure.

It is not necessarily token-internal grammar.

The solution is stratification.

Measure the effect within hands and within quires.

Does the Core Predict More Than the Suffix?

This is where the result can become far more informative.

If core identity predicts:

  • suffix;
  • line position;
  • next-token class;
  • section distribution;
  • label versus prose use;

then the core is behaving like a genuinely organising unit.

If its only success is predicting a suffix defined from the same decomposition, the interpretation remains narrower.

Independent downstream predictions are the stronger test.

The Directional Question

Does core predict suffix better than suffix predicts core?

If so, the asymmetry may reveal generative direction.

But prediction asymmetry can also arise from vocabulary size.

A small suffix inventory is easier to predict from many cores than a large core inventory is to predict from one suffix.

Therefore directional causal claims require entropy-normalised comparison, not raw accuracy alone.

This connects with The Directional Dissociation Problem.

What Would Make the Result Much Stronger?

  • Independent reimplementation of the decomposition.
  • Held-out quire validation.
  • Similar predictive strength across major transcription variants.
  • Robustness after uncertain spaces are merged or split.
  • Persistence within Currier and hand strata.
  • Additional predictions beyond suffix identity.
  • Successful reproduction by a historically plausible mechanism using fixed parameters.

What Would Weaken It?

  • Prediction collapses under alternative segmentation.
  • Most accuracy comes from one globally dominant suffix.
  • The effect disappears within scribal or Currier strata.
  • Core and suffix definitions leak information from the test set.
  • A simpler positional baseline performs equally well.

What the Core–Suffix Dependency Does Not Prove

  • It does not prove the core is a lexical root.
  • It does not prove the suffix is grammatical.
  • It does not prove the text is pronounceable.
  • It does not prove a cipher.
  • It does not prove shorthand.
  • It does not identify semantics.
  • It does not establish the historical direction of token construction.

What It Can Give Us

  • A quantified internal dependency target.
  • A way to compare morphology, cipher, shorthand and generator models.
  • A bridge from descriptive token regions to predictive testing.
  • A measurable constraint any future token generator must reproduce.

Primary School: The Middle Chooses the Ending

Create invented words where every middle type has only one legal ending.

Students quickly discover they can predict the ending before seeing it.

That is conditional structure.

Secondary School: Compare Against the Majority Baseline

Give students a dataset with three suffix classes.

First predict using the most common suffix only.

Then predict using core identity.

The improvement—not the final percentage alone—shows how much information the core contributes.

JC and Adult Readers: Freeze the Segmentation

Learn token regions on one training partition.

Freeze them.

Estimate H(suffix|core) on held-out quires.

Then compare:

  • global suffix entropy;
  • core-conditioned suffix entropy;
  • Currier-conditioned entropy;
  • hand-conditioned entropy;
  • core+Currier and core+hand models.

This identifies whether the core contributes information beyond manuscript strata.

Reader Checklist

  1. How is the token segmented?
  2. Were the regions defined prospectively?
  3. What is the suffix majority baseline?
  4. Is accuracy held out?
  5. Does the result survive Currier separation?
  6. Does it survive hand and quire separation?
  7. Does it survive alternate transcription?
  8. Does core identity predict anything else independently?
  9. Can a competing mechanism reproduce the same dependency?
  10. Is “suffix” being treated as a grammatical conclusion rather than an operational label?

Research Foundations

The Final Idea

The middle of a Voynich token may know something about how the token must end.

That is already significant.

What it knows—grammar, cipher state, shorthand class, notation rule or generator slot—is the question still open.

Prediction is evidence of structure. Naming the structure requires another layer of proof.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading