VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Edge Problem: Why More Structure May Live at Token Boundaries Than Between Whole Tokens

We have spent a century staring at Voynich words.

Perhaps some of the strongest order is sitting between them.

A modern transcription gives us convenient objects.

A glyph.

A token.

A space.

Then the machinery of language analysis starts almost automatically.

Glyph becomes letter.

Token becomes word.

Space becomes word boundary.

Those assumptions are useful because without them we cannot calculate much of anything.

They are dangerous because useful representation can become invisible theory.

In August 2026, Liudmila Rozanova and Alexander Temerev published a preprint designed to attack exactly those assumptions. Using the Zandbergen-Landini transliteration, matched prose, cipher and pseudo-text controls, and quire-level resampling, they report a peculiar structure:

the identity of one whole Voynich token predicts the next token astonishingly weakly.

But the glyphs sitting at the two edges of a token boundary remain statistically coupled.

At the same time, spaces already marked uncertain by transcribers behave differently from confident spaces and are physically narrower in the authors’ image-coordinate analysis.

This is the Edge Problem.

What if the visible “word” is not the level where Voynich sequence order is strongest—and what if the blank between two tokens is not always the kind of boundary we have assumed?

Quick Read

  • A 2026 arXiv preprint by Rozanova and Temerev tests three common assumptions: glyph = letter, token = word, blank = word space.
  • Under their analysis, exact Voynich token identity predicts the next token by under 1% of token entropy, lower than their matched prose controls, which fall around 2–10%.
  • Yet the glyphs immediately flanking token boundaries share about 0.2 bits of mutual information, greater than in their prose controls.
  • This produces a striking asymmetry: weak order among whole token identities, but stronger order at the token edges.
  • The authors also split spaces into confident and uncertain classes already present in the transliteration evidence.
  • Uncertain spaces behave more like word-internal junctures than confident separators in several of their tests.
  • They report that uncertain gaps are physically narrower, with an AUC of 0.905 in an independent image-coordinate analysis and the same direction in a small blind ink audit.
  • Learned multi-symbol units cross uncertain spaces even when spaces are erased before unit learning, suggesting that some transcription boundaries may split larger recurrent structures.
  • A self-citation generator and a Voynich-imitating cipher reproduce several other Voynich properties in the paper but fail the reported edge-glyph coupling and the manuscript’s very open, singleton-rich vocabulary.
  • The work is a 2026 preprint. It is important new evidence, not settled consensus.
  • The result does not prove that Voynich has no words. It shows that wordhood and boundary strength must be demonstrated rather than assumed uniformly.

The Strange Result: The Whole Token Knows Little, the Edge Knows More

Suppose ordinary Voynich tokens are words in the usual linguistic sense.

Then one natural place to search for syntax is:

word A → word B

The manuscript does show preferences of this kind. Our existing Syntax Before Semantics article owns that evidence.

But exact token identities are a very high-dimensional vocabulary.

If many Voynich tokens occur once, twice or only locally, exact token-to-token prediction will naturally be sparse.

Rozanova and Temerev therefore ask how much information the identity of one token carries about the identity of the next.

Their reported answer is surprisingly small: under one percent of token entropy.

Then they zoom inward by one level.

Ignore the complete words.

Look only at the last glyph before a space and the first glyph after it.

Those edge glyphs are coupled more strongly than comparable prose controls in the study.

That is not the pattern a naïve word model would make us expect.

The boundary may preserve a relation that disappears when we summarise each side as one whole token identity.

Why Exact Token Identity Can Be the Wrong Resolution

Imagine two ordinary English sentences:

the dogs are running

those cats were sleeping

If we treat every exact word as unrelated to every other exact word, much of the grammatical structure disappears.

Dogs and cats are different tokens but occupy related roles.

Are and were are different tokens but share a grammatical class.

Running and sleeping share an ending pattern.

Voynich may create an even more severe version of this problem because its word families are enormous and many exact forms are rare.

A weak exact-token predictor therefore does not automatically imply:

  • no syntax;
  • no semantic relation;
  • independent words;
  • random order.

It may mean exact token identity is not the right abstraction.

The edge result makes that possibility concrete.

An Edge Can Carry Grammar Without the Whole Word Doing So

Natural languages often place useful information at edges.

Prefixes.

Suffixes.

Clitics.

Agreement markers.

Sound changes across word boundaries.

A cipher can do the same for completely different reasons.

The final visible glyph of one group may encode a state that constrains the first glyph of the next.

An abbreviation system may make final marks signal the grammatical or scribal class of what follows.

A generated notation can deliberately choose compatible edge forms while allowing the interiors to vary.

So edge coupling is not a language detector.

It is a mechanism constraint.

The Space Problem Is Not Binary

Our original Spaces and Word Boundaries article asks whether Voynich “words” are really words.

The 2026 upgrade makes the problem sharper.

Not every visible gap is equally convincing.

Transcribers already mark some spaces as uncertain because the physical distance between neighbouring glyph groups is ambiguous.

That means the data already contains a graded observation:

  • confident separation;
  • uncertain separation.

The new preprint asks whether those classes behave differently statistically.

They report that they do.

Uncertain spaces behave more like internal junctures than confident word boundaries.

This is important because it connects a human palaeographic judgement to an independent distributional difference.

The Gap Is Physically Graded Too

The paper does not rely only on transliteration labels.

It also examines physical gap width using independent image coordinates.

The authors report that uncertain spaces are physically narrower than confident spaces, with an AUC of 0.905 for separating the classes by measured gap width.

They further report the same direction in a small blind ink audit.

Again, this is a preprint result.

But conceptually it is strong because two evidence streams agree:

  • transcribers find some gaps uncertain;
  • those gaps are, on average, physically narrower in the reported measurement.

The blank itself becomes measurable data.

Voynich may not have one universal space. It may have a continuum of separation whose endpoints our transcription converts into categories.

A Boundary Can Be Real Without Being a Word Boundary

This distinction matters enormously.

A scribe can place a small gap between:

  • two morphemes;
  • a stem and abbreviation;
  • two parts of a compound;
  • two cipher groups;
  • two graphical chunks of one word;
  • one word and the next word.

The physical gap is real.

Its linguistic status is not automatically known.

That is why saying “spaces are not spaces” is too theatrical if taken literally.

The paper’s stronger lesson is subtler:

blank width and functional boundary strength may not map one-to-one onto our binary transcription.

Learned Units Cross the Uncertain Gaps

Another experiment in the preprint removes all spaces before learning recurrent multi-symbol units.

If the original uncertain spaces represented strong word boundaries, we might expect learned units to avoid crossing them.

The authors report the opposite tendency: learned units cross uncertain spaces in a way compatible with those gaps being weaker internal junctures.

This matters because the result is not based on telling the learner which gaps were uncertain.

The spaces have been erased first.

The sequence itself recovers structures that cross those locations.

If replicated, this would make uncertain spaces one of the manuscript’s best examples of transcription metadata predicting independent structure.

The Edge Problem and Directional Dissociation Are Related—but Not the Same

The new Directional Dissociation Problem reports a split between token-internal directionality and boundary-level directionality.

The Edge Problem asks where the strongest boundary-scale information sits.

One result says:

different structural levels favour different directions.

The other says:

whole-token identity is a weak predictor, while boundary glyphs remain coupled.

Together they create a demanding architectural picture.

The visible token may be a container whose interior construction and boundary interaction are governed differently.

That is not a decipherment.

It is a more precise problem statement.

Why Self-Citation Passes Some Tests and Fails the Edge Test

The 2026 preprint includes a self-citation generator as one control.

This is useful because the generator has a known cause.

It can reproduce several Voynich-like behaviours:

  • low entropy;
  • recurrent multi-symbol unit scale;
  • weak whole-token predictive order;
  • resistance to a calibrated simple-substitution attack.

Yet the authors report that it does not reproduce the observed edge-glyph coupling.

That is exactly what a good discriminator should do.

It separates two corpora that look similar under older tests.

See The Self-Citation Problem.

Why a Voynich-Like Cipher Can Also Fail at the Edge

The same paper reports that a published Voynich-imitating cipher can reproduce several headline properties yet still fail the edge-glyph coupling criterion.

This is important because “cipher” is not one mechanism.

A cipher that successfully matches entropy and unit scale may still place the wrong information at token boundaries.

The edge therefore becomes a new place where competing cipher designs can be tested.

A historical mechanism should explain not merely why tokens look Voynich-like but why their entrances and exits relate the way they do.

Could This Be Sandhi-Like?

In natural languages, word boundaries can alter sound or form.

One word’s ending interacts with the next word’s beginning.

Historical spelling can sometimes reflect such interactions.

This makes a sandhi-like linguistic comparison reasonable.

But the analogy does not prove phonology.

The same edge interaction can arise in:

  • cipher group construction;
  • abbreviation conventions;
  • scribal joining rules;
  • generated chunk compatibility;
  • morphological agreement.

“Sandhi-like” is a useful comparator for the shape of the dependency.

It is not a recovered spoken language process.

Could the Edge Be an Encoding Checksum?

A constructed or cipher-like system can make token endings constrain token beginnings for procedural reasons.

An ending can select a table.

A beginning can encode a class.

Adjacent groups can be required to avoid repeated patterns.

A boundary form can carry state from one group into the next.

This makes edge coupling particularly interesting to cryptanalytic models.

But a checksum-like or state-transfer hypothesis must produce an executable rule.

Which ending constrains which beginning?

How?

Does the rule vary by Currier regime?

Does it survive line boundaries?

A mechanism earns seriousness through predictions, not through technical vocabulary.

Could We Have Cut One Larger Unit Into Two Tokens?

This is perhaps the simplest explanation of edge coupling.

Suppose one real structural unit crosses a weak visible gap.

The transcription splits it into token A and token B.

Now the last glyph of A and first glyph of B appear unusually dependent.

Of course they are.

They were originally internal neighbours inside one larger unit.

This explanation predicts a strong difference between confident and uncertain spaces.

That is why the reported physical gap-width result is so relevant.

If edge coupling concentrates at uncertain/narrow gaps, segmentation error becomes a serious candidate.

If it persists equally across confident wide spaces, a true cross-boundary rule becomes more plausible.

A Better Boundary Model Is Graded

Instead of one binary field:

space / no space

we may need a richer representation.

  • no visible gap;
  • very weak gap;
  • uncertain gap;
  • confident gap;
  • line boundary;
  • paragraph boundary;
  • diagram/locus boundary.

Each boundary type can then be tested separately.

Does edge coupling strengthen or weaken with physical gap width?

Does a line break behave like a wide space?

Does a paragraph break reset the relation?

Do labels have the same edge grammar as prose?

The manuscript may tell us which boundaries are functionally equivalent if we stop forcing them into one category first.

Why This Matters for the Word-Length Problem

If weak spaces split larger units, observed token lengths are partly a product of transcription boundary decisions.

Merge uncertain splits and the token-length distribution changes.

Some extremely short tokens disappear.

Some common lengths shift upward.

The narrow Voynich length envelope may remain.

But its exact shape depends on the segmentation.

This is why the Word-Length Problem and Edge Problem must stay linked.

Neither should borrow the other’s assumptions invisibly.

Why This Matters for Syntax Before Semantics

Whole-token next-word statistics remain useful.

They may simply be too coarse by themselves.

Suppose two large word families behave similarly but exact token forms vary widely.

Exact token A may predict exact token B poorly.

But suffix-class A may strongly predict prefix-class B.

That would look like weak word syntax and strong edge syntax.

A better model should therefore compare multiple resolutions:

  • exact token;
  • word family;
  • prefix class;
  • suffix class;
  • last glyph;
  • first glyph;
  • learned multi-symbol unit.

The Hapax Problem Makes Whole-Token Models Especially Fragile

The same 2026 preprint reports an unusually open vocabulary with a high proportion of singleton token types—around seventy percent under the study’s counting choices.

A vocabulary dominated by forms seen once cannot support rich exact-token transition estimates.

The data is too sparse.

But those singleton forms can still share:

  • prefixes;
  • suffixes;
  • edge glyphs;
  • internal units;
  • positional templates.

This is another reason the edge level can carry recoverable structure when whole-token identity does not.

Transcriber Disagreement Becomes Evidence Rather Than Annoyance

For decades, uncertain spaces have often been treated as a nuisance to be resolved.

Choose one transcription.

Move on.

The graded-boundary result suggests a better attitude.

Disagreement can be data.

If several expert transcribers hesitate at the same locations, the physical manuscript may itself contain a weaker boundary class there.

Instead of collapsing the uncertainty, preserve it.

Then ask whether the text behaves differently around high-uncertainty boundaries.

This is a broader lesson for manuscript science:

uncertainty is sometimes a property of the evidence, not a defect in the researcher.

Replication Has to Touch the Images

This result is too important to remain only inside a transliteration file.

A strong independent replication should return to the scans.

  • Measure physical gap widths independently.
  • Blind the measurer to transcription confidence labels.
  • Use several folios and scribal regimes.
  • Control for glyph width and neighbouring shapes.
  • Correct image scale and distortion.
  • Compare existing transliterators’ uncertainty marks.
  • Test whether learned units cross the same weak gaps under different algorithms.

If independent image measurements reproduce the graded-space effect, the result becomes substantially harder to dismiss as transcription convention.

Replication Has to Survive Currier, Scribe and Quire

A manuscript-wide average can hide multiple causes.

Edge coupling might be:

  • strong in Currier A and weak in B;
  • specific to one proposed hand;
  • strong on text pages and weak on labels;
  • driven by one quire;
  • different around uncertain spaces.

The 2026 paper uses quire-level resampling, which is a valuable start because it reduces the risk that one local cluster dominates the result.

Future work should go further and publish the effect size by regime.

The Strongest Possible Outcome Is Not “Words Are Wrong”

Science often gets weakened by dramatic phrasing.

Suppose the edge findings replicate perfectly.

Would that prove Voynich tokens are not words?

No.

It would prove something more precise:

our usual exact-token and binary-space representation does not capture all the manuscript’s sequential structure efficiently.

Some tokens may be words.

Some boundaries may be genuine word spaces.

Some uncertain gaps may be internal.

Some edge glyphs may carry morphology or control state.

The correct future model can be mixed.

What the Edge Problem Does Not Prove

  • It does not prove Voynich has no words.
  • It does not prove all spaces are meaningless.
  • It does not prove uncertain spaces are always internal.
  • It does not prove the text is a cipher.
  • It does not prove a spoken sandhi process.
  • It does not prove self-citation is wrong in every form.
  • It does not identify plaintext units.
  • It does not establish that the 2026 preprint will survive all independent replication.

What It Can Give Us

  • A new location for Voynich order: token edges.
  • A measurable distinction between whole-token order and edge-level order.
  • A graded rather than binary boundary model.
  • A bridge between transcription uncertainty and physical gap width.
  • A discriminator that separates Voynich from some strong synthetic controls.
  • A reason to model learned multi-symbol units across uncertain spaces.
  • A stronger test for every future language, cipher, abbreviation or generator hypothesis.

Primary School: The Doorway Can Have a Rule

Imagine two rooms full of mixed toys.

The toys inside each room seem almost random.

But every door follows a rule:

a red toy must be nearest the door on the left side and a blue toy nearest the door on the right.

If you study whole rooms, the rule is hard to see.

If you study doorways, it becomes obvious.

That is the Edge Problem.

Secondary School: Merge the Weak Spaces

Take a synthetic text with some spaces deliberately made narrower than others.

Transcribe all gaps as equal spaces.

Measure word length and next-word statistics.

Then merge the words across the narrowest gaps and measure again.

Students see how a tiny segmentation decision can change the statistical object.

JC and Adult Readers: Model the Boundary Strength Explicitly

For every candidate token boundary, preserve:

  • physical gap width;
  • transcriber confidence;
  • left-edge glyph;
  • right-edge glyph;
  • line position;
  • Currier regime;
  • hand/quire;
  • whether a learned unit crosses it.

Then model boundary strength as a variable rather than forcing every gap to equal 1.

The question becomes:

where does the evidence itself say one unit ends?

Reader Checklist: Before You Treat Every Voynich Space as a Word Boundary

  1. Is the gap marked confident or uncertain?
  2. How wide is it physically?
  3. Does the same transcriber treat comparable gaps consistently?
  4. Do other transcriptions agree?
  5. Do recurrent learned units cross it?
  6. Is edge-glyph coupling stronger at that boundary class?
  7. Does merging uncertain gaps change word-length conclusions?
  8. Does it change next-token conclusions?
  9. Does the effect survive quire-level resampling?
  10. Does it survive Currier A/B separation?
  11. Do known cipher and generator controls reproduce it?
  12. Is the article claiming “not a word” when the evidence only supports “boundary strength uncertain”?

Frequently Asked Questions

What is the Voynich Edge Problem?

It is the reported 2026 mismatch in which exact whole-token identity predicts the next token weakly, while glyphs immediately flanking token boundaries retain stronger statistical dependence.

Does that mean Voynich tokens are not words?

Not necessarily. It means exact token identity may be too coarse to capture the strongest sequential structure and that some spaces may represent weaker or different kinds of boundaries.

What are uncertain spaces?

They are gaps whose status as true token boundaries is visually ambiguous in the manuscript and is represented as uncertain in transcription. The 2026 preprint reports that they are physically narrower and behave more like internal junctures than confident spaces.

Is this peer reviewed?

The Rozanova–Temerev study discussed here is an August 2026 arXiv preprint. Its measurements are important and unusually recent, but independent replication and review remain necessary.

Why is edge coupling useful?

Because several sophisticated Voynich-like controls reproduce older headline statistics yet fail this reported boundary signature. It therefore offers a new discriminatory test for competing mechanisms.

Research Foundations

The Final Idea

We gave Voynich a vocabulary because the spaces made one easy to see.

That was sensible.

Now the manuscript may be asking us to look one resolution lower.

At the last glyph.

The first glyph.

The width of the blank between them.

The unit that may cross that blank.

Perhaps the space between two Voynich tokens is not empty. Perhaps it is one of the places where the system tells us what kind of units we should have been looking for all along.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading