VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Long-Range Memory Problem: What Survives Beyond the Next Word?

The Voynich Manuscript is famous for what happens next door.

Similar words cluster locally.

Characters obey strong internal constraints.

Line beginnings behave differently from line endings.

But there is a more difficult question.

What happens fifty characters later?

A paragraph later?

A page later?

Does the manuscript carry statistical memory beyond the local neighbourhood?

For more than two decades, researchers have reported forms of long-range structure in Voynichese. The methods differ. The units differ. The conclusions differ.

That disagreement is not a nuisance.

It tells us that “long-range correlation” is not one phenomenon.

The real problem is not whether Voynich has memory. It is identifying what kind of memory survives, at what scale, and which mechanism could have created it.

Quick Read

  • Voynich researchers have reported non-local statistical structure at character, token and vocabulary-distribution scales.
  • Gabriel Landini’s early computational work and later studies noted long-range correlations unusual enough to deserve comparison with meaningful text and generated text.
  • Montemurro and Zanette’s 2013 PLOS ONE study found long-range word-distribution structure and information-theoretic “keywords” compatible with topic organisation.
  • A 2021 PLOS ONE study of symbol roles reported unusual autocorrelation patterns across many individual glyphs.
  • Recent 2026 community experiments use character-level mutual information and shuffle tests to ask where the long-range signal lives.
  • In those experiments, the original sequence can retain more long-range information than token-shuffled or character-shuffled controls.
  • Line shuffling often preserves more structure than token shuffling, suggesting some dependence lives inside lines or local page organisation.
  • However, similar shuffle-sensitive long-range behaviour can also arise in known recursive generators such as Timm’s self-citation text.
  • Therefore long-range structure does not by itself prove natural language or semantic prose.
  • Possible sources include topic clustering, document organisation, copying/generation memory, cipher state, scribal habits, line-level production and real semantic sequence.
  • The strongest future test separates these sources by using matched controls and by shuffling one structural level at a time.

What Does “Memory” Mean Here?

Not human memory.

Not necessarily semantic memory.

Statistical memory means that knowing something about the sequence at one position changes what we can predict at a later position.

In a completely shuffled character stream, distant positions become nearly independent.

In structured text, dependencies can persist.

  • A topic can keep certain vocabulary active for several paragraphs.
  • A grammatical construction can span several words.
  • A recursive generator can keep reusing descendants of earlier forms.
  • A cipher can preserve state across several encoded groups.
  • A scribe can maintain one local spelling convention for a page or quire.

All of these create memory.

They do not create the same memory.

The Old Observation: Voynich Is Not a Bag of Independent Tokens

The idea that Voynich contains non-local structure is not new.

Early computational studies by Gabriel Landini and others examined serial correlation and reported behaviour that did not collapse into simple random text.

René Zandbergen’s research synthesis preserves several of these older lines of work under the broader heading of long-range correlations.

The methods are historically important but should not be treated as one modern unified statistic.

Their common lesson is more robust:

where a Voynich symbol or word appears can depend on more than its immediate neighbour.

Montemurro and Zanette: Long Range at the Vocabulary Scale

In 2013, Marcelo Montemurro and Damián Zanette used information-theoretic methods to study word distributions across the manuscript.

Their central insight was spatial.

Some token types are distributed relatively evenly.

Others concentrate in particular regions.

That is what a topic-bearing text often does.

A chapter about plants uses one vocabulary.

A chapter about astronomy uses another.

The researchers extracted information-rich token networks and argued that the organisation was compatible with genuine meaningful text.

That finding remains important.

But “compatible with meaningful text” is not “meaning established”.

A generator whose parameters change by section could also create regional vocabulary.

A cipher whose tables change by quire could do the same.

The canonical local-vocabulary discussion remains Keywords Without Meanings.

Symbol Autocorrelation: Memory Below the Word

A 2021 PLOS ONE study revisited Voynich symbol roles using several statistical tools, including autocorrelation.

The authors reported unexpectedly persistent positive autocorrelations for many individual symbols over multiple lags.

This is interesting because it moves the memory question below whole-token semantics.

A character can recur according to a delayed rule even if we do not know whether the character is a letter.

Possible causes include:

  • repeated word families;
  • slot-like token construction;
  • local scribal state;
  • section-specific glyph preferences;
  • recursively generated vocabulary;
  • encoded plaintext structure.

Again, the signal is real enough to test.

The cause remains a competition among models.

The 2026 Shuffle Question: What Breaks When We Randomise Order?

Recent specialist community experiments have made the long-range question more intuitive by using destructive controls.

Take the Voynich character stream.

Measure mutual information at increasing distances.

Then alter the sequence in controlled ways.

  • Character shuffle: destroys nearly all ordering.
  • Token shuffle: preserves internal token spelling but destroys token order.
  • Line shuffle: preserves within-line ordering but destroys longer continuity among lines.
  • Original: preserves every observed sequence relation.

In reported 2026 experiments, the ordering typically appears:

Original > line shuffle > token shuffle > character shuffle

That pattern suggests that some long-range signal depends on the ordering of tokens and lines, not merely on internal token spelling.

These are community experiments rather than peer-reviewed consensus and should be treated accordingly.

But the experimental design is valuable because it tells us what kind of intervention a future formal study should reproduce.

Why Shuffling Is Such a Powerful Test

Shuffling preserves some things and destroys others.

If you shuffle whole tokens:

  • token frequencies remain identical;
  • token spellings remain identical;
  • word-length distribution remains identical;
  • internal character structure remains identical;
  • the sequence among tokens is destroyed.

Therefore any statistical difference between original and token-shuffled text must come from sequence relations above the individual-token level.

This is clean experimental logic.

It does not tell us what that sequence relation means.

The Important Complication: Self-Citation Has Long-Range Memory Too

A recursive generator provides a crucial control.

In Timm–Schinner self-citation, later tokens descend from earlier tokens.

The resulting corpus therefore contains a historical dependency chain.

Shuffle the words and you break that chain.

Long-range mutual information can change dramatically even though the generated text contains no ordinary semantic prose.

This demolishes another weak argument:

Voynich has long-range correlations, therefore it must contain meaningful natural language.

No.

Long-range dependence proves dependence.

Its cause must still be identified.

See The Self-Citation Problem.

Natural Language Memory and Generator Memory Can Look Different

This is where the research becomes more useful.

Natural language often carries long-range structure through topic, reference, grammar and discourse.

A recursive generator carries long-range structure through genealogy.

Those mechanisms respond differently to interventions.

For example, shuffling whole words in normal prose destroys grammar and sentence meaning but may preserve broad character-frequency tails generated by the vocabulary itself.

In a recursive generator, word order may directly encode the ancestry relation that created the vocabulary.

Some 2026 exploratory analyses report that the long-range MI response to word shuffling differs between natural-language corpora, Naibbe-like cipher output, Timm-generated text and Voynich.

Those comparisons are promising.

They are not yet a settled discriminator.

The Page and Quire Can Create Long-Range Structure Without Syntax

Suppose a whole quire uses a distinct scribal convention.

One glyph variant becomes common for fifty lines.

Then the next quire changes hand or text regime.

A character-level long-range statistic will detect persistence.

But the source is production state, not sentence grammar.

The same happens with topic.

If herbal pages favour one vocabulary family for several folios, distant tokens become correlated because they share document domain.

Therefore long-range analysis must control for:

  • Currier regime;
  • scribe/hand;
  • quire;
  • visual section;
  • line boundaries;
  • page boundaries.

The Line May Be a Memory Container

Voynich lines already behave as important units.

Initial and final positions differ.

Some forms cluster within lines.

If line shuffling preserves substantially more signal than token shuffling, one possibility is that the line itself contains a local dependency architecture.

That could arise from:

  • line-by-line composition;
  • line-local copying;
  • layout-conditioned token choice;
  • short-range syntax;
  • cipher state reset at line boundaries.

The existing Line as a Unit article owns the positional evidence. Long-range memory asks whether line units also carry sequential persistence.

The Difference Between Burstiness and Memory

A token can be bursty without one occurrence causing the next.

If a page topic activates a technical term, that term appears many times locally.

The pattern looks clustered.

A recursive generator can also create bursts because one form keeps spawning descendants nearby.

These causal histories differ.

One is shared topic state.

One is genealogical copying state.

A good long-range study should distinguish:

  • clustering of the same token;
  • clustering of related token families;
  • character-level dependence;
  • document-level topic persistence.

Why Whole-Manuscript Concatenation Is Dangerous

A corpus file often places every page one after another.

That creates artificial neighbours.

The last character of one folio becomes adjacent to the first character of the next even if the physical reading sequence is uncertain.

Foldouts are linearised.

Missing leaves disappear.

Current foliation can masquerade as original composition order.

Therefore long-range analysis should be repeated under several sequence regimes:

  • within-line only;
  • within-page only;
  • within-bifolium;
  • within-quire;
  • current manuscript order;
  • candidate reconstructed orders where justified.

Any effect that survives all of these becomes much more interesting.

Long-Range Memory and the Drift Problem

Slow change across the manuscript can mimic memory.

If one token family gradually rises while another declines, nearby pages become more similar than distant pages.

This creates long-range structure even without direct token-to-token causation.

The Drift Problem therefore becomes an essential control.

Ask whether the apparent memory remains after slow quire/page trends are removed.

If it does, a more local sequential mechanism becomes more plausible.

Long-Range Memory and the Exemplar Problem

Copying from an exemplar can also create persistence.

Suppose several consecutive pages are copied from one source section.

The source carries its own vocabulary and formatting.

When the source changes, the statistical state changes.

The surviving Voynich then contains long-range clusters reflecting source organisation rather than one continuous generated stream.

See The Exemplar Problem.

What a Serious Long-Range Test Should Do

  1. Declare the transcription and glyph segmentation.
  2. Declare whether spaces are retained.
  3. Keep line/page/quire boundaries explicit.
  4. Measure original long-range dependence.
  5. Shuffle characters while preserving frequencies.
  6. Shuffle tokens while preserving token spellings.
  7. Shuffle lines while preserving line-internal order.
  8. Shuffle pages within controlled sections.
  9. Compare natural-language controls matched for length and genre.
  10. Compare known cipher controls.
  11. Compare known generators, especially self-citation.
  12. Separate Currier A/B and scribal groups.
  13. Use held-out quires to test any proposed mechanism.

What Long-Range Memory Does Not Prove

  • It does not prove natural language.
  • It does not prove semantic prose.
  • It does not prove self-citation.
  • It does not prove a cipher.
  • It does not prove current page order is original.
  • It does not prove one mechanism operates at every scale.
  • It does not make every mutual-information curve a direct measure of “meaning”.

What It Can Give Us

  • Evidence that some structure persists beyond immediate token interiors.
  • A way to locate which organisational scale carries the persistence.
  • A discriminator among independent-token, line-local, recursive and discourse-like models.
  • A bridge between local word families and manuscript-wide drift.
  • A destructive test: shuffle one layer and observe what information disappears.

Primary School: Shuffle the Story

Write six sentences about one trip to the zoo.

Then shuffle every word.

The same words remain.

The story disappears.

Now shuffle only the six sentences.

Much more information survives.

The exercise shows how changing order at different levels destroys different kinds of structure.

Secondary School: Four Shuffles

Compare a text under four conditions:

  • original;
  • line shuffle;
  • word shuffle;
  • character shuffle.

Ask which patterns survive each intervention.

This is the experimental logic behind long-range Voynich analysis.

JC and Adult Readers: Decompose Memory by Scale

Measure the same statistic after progressively destroying structure.

Then attribute only the lost component to the layer that was shuffled.

If token shuffle removes most of the signal, token order matters.

If line shuffle removes little, cross-line order contributes little.

If page shuffle removes a large remainder, page sequence or page-state persistence matters.

This converts “Voynich has correlations” into a map of where correlation lives.

Reader Checklist: Before You Call a Voynich Correlation “Meaning”

  1. What unit is correlated—character, token, line, page or vocabulary?
  2. At what distance?
  3. What null model is used?
  4. Does token shuffling remove the effect?
  5. Does line shuffling remove it?
  6. Does page/quire structure explain it?
  7. Could Currier drift create the same pattern?
  8. Could scribal hand persistence create it?
  9. Does self-citation reproduce it?
  10. Does a known cipher reproduce it?
  11. Does the result survive alternate transcription?
  12. Is “correlation” being silently translated into “semantics”?

Frequently Asked Questions

Does Voynich have long-range correlations?

Multiple studies using different methods have reported non-local structure. The precise source and interpretation of that structure remain debated.

Does that prove the text has meaning?

No. Meaningful language can create long-range dependence, but so can recursive generators, topic-conditioned pseudo-text, cipher state and production drift.

Why shuffle the text?

Shuffling destroys selected kinds of order while preserving others, allowing researchers to identify which structural level contributes to a measured signal.

Are the 2026 shuffle results peer reviewed?

The specific recent mutual-information shuffle experiments discussed here are exploratory specialist-community work. They are useful as test designs and emerging evidence, not settled consensus.

Research Foundations

The Final Idea

A manuscript can remember without remembering meaning.

A page can remember its topic.

A generator can remember its source words.

A cipher can remember its state.

A scribe can remember a local convention.

And a meaningful text can remember what it was talking about.

The next question is not whether Voynich has long-range structure. It is which kind of history that structure is carrying forward.


Continue Through the Voynich Research Map

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading