VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | Exact Repetition: Why Voynich Repeats Words but Rarely Repeats Phrases

The Voynich Manuscript has a repetition paradox.

It can repeat a word-like unit shamelessly.

The same token appears twice.

Then three related or identical-looking forms can sit beside one another.

A famous line in the human-figure section contains a run in which essentially the same long token occurs repeatedly with only small variation.

At first glance, this looks like a highly repetitive text.

Then researchers count exact two-word and three-word sequences.

The impression changes.

Voynich has surprisingly few repeated exact word bigrams and trigrams given how often its individual tokens recur.

In a 2011 computational review, Sravana Reddy and Kevin Knight described the manuscript’s whole-token word order as unusually weak: knowing the previous token improved prediction of the next token only marginally compared with several natural-language controls.

So we have two truths at once.

  • Voynich readily repeats individual tokens and near-identical local forms.
  • Voynich is comparatively reluctant to repeat exact longer token sequences.

That combination is more interesting than either fact alone.

A language can repeat words.

A list can repeat fields.

A cipher can suppress repeated phrases.

A generator can deliberately mutate repeated forms so that exact phrases almost never recur.

A copying workflow can produce local echoes without creating reusable syntax.

The repetition problem is not “Why does Voynich repeat?” It is “Why does repetition remain local while exact phrase recurrence stays weak?”

That question belongs between word families and syntax.


Quick Read

One-sentence answer: Voynich text contains conspicuous exact token doubling and short local runs, but exact multi-token phrases recur far less strongly than expected from ordinary language controls, creating a distinctive architecture in which repetition is concentrated at the level of individual forms, local variants and line-scale production rather than stable repeated phrases.

  • Exact token repetition occurs in the Voynich Manuscript.
  • Repeated and near-repeated forms can cluster closely on a line.
  • This article distinguishes exact repetition from the broader near-neighbour family phenomenon covered separately.
  • Reddy and Knight’s 2011 analysis reported very few repeated word bigrams or trigrams despite ordinary-looking unigram word entropy.
  • In their VMS-B sample, using the previous whole token improved next-token prediction only modestly compared with English, Arabic, Chinese and Hungarian controls.
  • This suggests that exact token identity carries surprisingly little information about the identity of the next whole token.
  • Yet the final characters of one token can carry information about the first one or two characters of the next token, so cross-boundary dependence is not absent.
  • Repeated local runs often respect line boundaries, linking repetition to the manuscript’s line architecture.
  • Natural language can contain doubling for emphasis, lists, formulae, names, quantities or rhetorical structure.
  • A record or tabular system can repeat identical fields without repeating full sentences.
  • An encoding system can map repeated plaintext phrases to varying ciphertext and therefore suppress exact phrase recurrence.
  • A local-generation system can create repeated or near-repeated tokens while intentionally avoiding exact long sequences.
  • No single mechanism is established by repetition alone.
  • A strong theory must explain both the repeats and the missing repeats.

Negative evidence is central here.

The phrases Voynich does not repeat are part of the signal.


Exact Repetition Is Not the Same as Word-Family Similarity

This distinction keeps the series clean.

The Word Families article owns tokens that are almost alike.

One glyph changes.

A prefix appears.

An ending changes.

This article begins one step stricter.

When does the same token appear again?

Then it asks a second question.

When does the same sequence of tokens appear again?

The gap between those answers is the repetition paradox.


Doubled Words Are Real

Voynich contains cases in which one space-delimited token is followed immediately by the same token.

These are visually striking because modern prose usually avoids accidental exact duplication.

But repeated words are not impossible in ordinary language.

We say:

  • very, very;
  • no, no;
  • bye-bye;
  • had had;
  • that that.

Other languages use productive reduplication for grammatical or semantic purposes.

Lists can repeat identical category markers.

Recipes can repeat quantities or operations.

Therefore a doubled Voynich token is compatible with language.

The real question is whether doubling follows a stable rule.

Which tokens double?

Where?

At what line position?

Does doubling change by Currier regime?

Those distributions can turn repetition from curiosity into grammar.


Runs of Repeated Forms Are More Striking Than Simple Doubling

Some Voynich passages place several identical or almost-identical long tokens side by side.

Reddy and Knight highlighted one such line in their 2011 review because it makes the local repetition phenomenon impossible to miss.

The sequence contains repeated instances of the same long token family with exact duplicates embedded among near-duplicates.

This can arise from several very different mechanisms.

  • linguistic reduplication or rhetorical repetition;
  • a list in which one field repeats;
  • copying from immediately adjacent text;
  • local generation by duplication and mutation;
  • encoded repetitions whose surface forms remain partly related.

The run is therefore a stress test.

A theory should explain why the same token can occur several times without letting that freedom explode into frequent repeated phrases everywhere else.


Why Repeated Bigrams Matter

A bigram at the word level is simply two consecutive tokens.

Known language repeats common two-word sequences constantly.

English repeats:

  • of the;
  • in the;
  • to the;
  • it is;
  • there are.

Technical prose repeats formulaic pairs even more strongly.

Recipes repeat “take X”.

Registers repeat field sequences.

So when a corpus has many recurring individual words but relatively few exact recurring word pairs, something unusual is happening at the transition level.

The token vocabulary repeats.

The token order does not repeat proportionally.

This is one reason Voynich syntax remains difficult to model as ordinary phrase structure.


Trigrams Raise the Standard Again

Three-token sequences are harder to repeat by accident.

In ordinary prose, frequently used constructions create recurring trigrams.

In formulaic technical writing, this can be even stronger.

Voynich’s scarcity of repeated exact bigrams and trigrams is therefore especially notable if one assumes the spaces delimit ordinary lexical words and the paragraphs are ordinary continuous prose.

That assumption may be wrong.

Perhaps spaces divide units below or above words.

Perhaps the text is encoded in a way that diversifies repeated phrases.

Perhaps the generative process intentionally constructs each token independently enough that phrase recurrence remains weak.

Phrase scarcity therefore feeds directly back into the word-boundary problem.


Whole-Word Prediction Is Surprisingly Weak

Reddy and Knight quantified part of this issue by asking how much knowing the previous token improves prediction of the next.

In their Currier-B sample, whole-word bigram context improved prediction only modestly over unigram frequency.

The improvement was much smaller than in English, Arabic or Hungarian controls and smaller than in their Chinese comparison.

This does not mean “there is no syntax”.

It means exact previous-token identity is unusually weak as a predictor of exact next-token identity.

That distinction matters because syntax can operate over categories rather than exact words.

“The” predicts a noun class, not one particular noun.

If Voynich has latent classes, whole-word bigrams can look weak while class-level grammar remains strong.

The fourth article in this batch owns that next level.


Yet Token Edges Do Talk to Each Other

This is where the repetition story becomes more subtle.

Reddy and Knight also examined whether the final characters of one token help predict the first characters of the next.

They found that while whole-word prediction remained weak, short edge-level context could improve prediction of the first one or two characters of the next token.

This echoes older Currier observations that word-final characters can affect the initial character of the following word-like unit.

So the boundary is not empty.

The manuscript may have:

  • weak exact word-to-word identity constraints;
  • stronger class or edge-to-edge constraints.

This is an important clue about where syntax-like order may live.

Perhaps the correct units are not full EVA tokens.

Perhaps endings and beginnings carry transition information.


Exact Repetition Often Feels Line-Bounded

The line article established that Voynich writing responds strongly to physical line boundaries.

Repetition adds another layer.

Currier noted that repeated sequences do not simply continue across line breaks as though the line were arbitrary wrapping.

Local repeated runs tend to live within the production space of a line.

This makes at least three models plausible.

  • The line is a semantic or record unit.
  • The line is a production buffer within which copying or generation occurs.
  • An encoding state resets at the line boundary.

All three can make repetition line-local.

The key is whether repeated tokens carry stable meanings or functions across different lines.


Natural Language Can Repeat Without Repeating Phrases Much

Weak exact phrase recurrence does not automatically exclude language.

A language with rich morphology can express the same construction with changing surface forms.

A flexible word-order language can vary sequence.

A large vocabulary can reduce exact phrase repeats.

Technical labels and names can create many low-frequency tokens.

But Reddy and Knight compared Voynich with Hungarian specifically because Hungarian provides a natural-language example with relatively flexible word order.

The Hungarian control still showed much stronger improvement from previous-word context.

This does not settle language identity.

It raises the burden on simple “ordinary prose with unusual spelling” models.

A language theory needs to explain why Voynich exact word-order recurrence is so weak.


A List or Record System Can Repeat Fields Instead of Phrases

Suppose a line is not a sentence.

It is a record.

Each record contains fields such as:

  • identifier;
  • class;
  • quantity;
  • condition;
  • operation.

Fields can repeat individually while exact full records remain rare.

This architecture naturally produces recurring tokens without recurring bigrams or trigrams at ordinary-language rates.

Quire 20 makes this possibility especially interesting because its star-marked entries are explicitly segmented into short record-like units.

If repetition patterns differ between Quire 20 records and long prose-like paragraphs, genre becomes a discriminator.

The manuscript may use several textual grammars for several jobs.


A Cipher Can Suppress Exact Phrase Repetition

One historical reason for using more complex ciphers is precisely to avoid obvious repetition.

If the same plaintext word can be represented by several ciphertext groups, repeated plaintext phrases need not create repeated visible phrases.

Homophonic systems distribute one plaintext unit over multiple ciphertext values.

Verbose systems can produce related but non-identical groups.

State-dependent systems can vary output by context.

This gives ciphertext a natural route to:

  • repeated underlying vocabulary;
  • weak exact visible phrase recurrence;
  • strong visible word families.

The 2025 Naibbe study is relevant because it demonstrates constructively that historically plausible hand ciphers can preserve meaningful plaintext while generating several Voynich-like statistical effects.

That does not prove Voynich repetition is cryptographic.

It makes phrase suppression a serious cipher prediction rather than a hand-wave.


Local Generation Can Produce the Paradox Naturally

A local-copying or mutation model has an obvious route to the observed pattern.

Select a nearby token.

Sometimes copy it exactly.

Usually alter one component.

Use the new token as another seed.

This naturally creates:

  • occasional exact doubling;
  • many near-repeats;
  • strong local families;
  • few exact repeated long phrases.

This is one reason Timm-style generation hypotheses treat local repetition as central evidence.

But as the Word Families article emphasised, the generator then has to explain the manuscript above token level.

Why different Currier states?

Why labels?

Why long-range topic-like distribution?

Why images?

Surface repetition is one layer of the manuscript, not the whole explanation.


Copying From a Formulaic Source Can Create Local Echoes Too

There is another alternative between ordinary language and on-the-fly generation.

The source itself may be repetitive.

A paradigm lists related grammatical forms.

A recipe table repeats ingredient fields.

A lexicon places related entries together.

A codebook groups similar codes.

A scribe copying such a source would produce local repetition without independently generating it.

This is why local echo is not proof of copy-and-modify generation.

The organisation may live in the exemplar.

Production model and source model must be separated.


Repeated Tokens Can Be Grammatical

If Voynich tokens correspond to linguistic units, some repeated forms may perform grammatical roles.

A duplicated item could mark:

  • plurality;
  • intensity;
  • continuation;
  • distribution;
  • iteration;
  • discourse emphasis.

Reduplication is a real linguistic phenomenon.

But a grammatical reduplication theory must be systematic.

The same type of token should double in similar syntactic environments and produce related semantic effects.

If doubling occurs indiscriminately across unrelated classes, a purely grammatical explanation weakens.

We need distribution before we need a translation.


Repeated Tokens Can Be Numeric or Quantitative

Another tempting interpretation is quantity.

Repeat one token twice to mean two units.

Repeat three times to mean three.

This is possible in specialised notation.

It should not be assumed.

A numeric-repetition model predicts:

  • controlled run lengths;
  • association with objects that can be counted;
  • consistent interaction with other quantity-like forms;
  • similar behaviour across page families if the notation is general.

Our retained research did not support a simple universal three-unit dosage code.

That negative result remains important.

Repeated forms may still encode quantity locally.

They are not established as a manuscript-wide dosage numeral system.


Repetition Can Be an Error—But “Error” Must Have a Distribution

Scribes duplicate words accidentally.

A copyist looks back to the exemplar and repeats the same unit.

This is called dittography in manuscript studies.

Could Voynich doubling simply be copying error?

Some cases may be.

But if repetition occurs systematically in certain contexts or with certain token classes, accident becomes less sufficient.

An error hypothesis should predict error-like behaviour.

Irregularity.

Correction sometimes.

No stable grammatical distribution.

“Scribal error” should explain residual cases.

It should not become the universal label for every repeat a theory cannot interpret.


The Scarcity of Repeated Phrases Is More Informative Than One Dramatic Run

A spectacular repeated line attracts attention because it is visible.

Corpus-wide absence is harder to see.

Yet the absence of many repeated exact bigrams and trigrams is arguably the more important observation.

It tells us that the dramatic run is exceptional inside a system that usually resists exact phrase reuse.

Any theory built from the run has to reproduce the background rarity.

If local generation explains the run, why are exact two-token repeats not much more common?

If formulaic language explains the run, why do formulae not recur more often elsewhere?

If a cipher suppresses phrase repetition, why does it sometimes allow exact doubling?

The exception and baseline must be explained by the same model.


Repetition Depends on What Counts as the Same Token

Transcription returns again.

Two visible forms may be treated as identical under a coarse alphabet and different under a fine one.

A compound gallows may be segmented differently.

An uncertain space can merge two tokens.

Therefore repetition counts depend on representation.

The strongest repetition phenomena should survive reasonable transliterations.

Where a result depends on one disputed glyph, image review should lower confidence.

A repeated phrase is only as exact as the transcription that defines exactness.


Repetition Depends on Currier Regime

A and B have different token inventories and family frequencies.

That changes the baseline probability of repetition.

A common B token has more opportunities to double on a B page than on an A page.

Therefore a raw repeat count across the whole manuscript can mix regimes unfairly.

A strong analysis should ask:

  • Does B repeat exact tokens more often after controlling for token frequency?
  • Do A and B differ in repeated phrase scarcity?
  • Are edge-to-edge dependencies different?
  • Do repeated runs correlate with one proposed hand?

Repetition may reveal another layer of the Currier distinction.

Or it may disappear after frequency is controlled.

Either outcome is informative.


Repetition Depends on Document Role

Labels should not be expected to repeat like prose.

Their vocabulary is flatter and many labels are unique.

Quire 20 records may repeat fields differently from long paragraphs.

Circular text may repeat category markers around a ring.

Pooling all text roles can therefore obscure genre-specific repetition.

A true recipe corpus might contain repeated operation phrases.

A label corpus might intentionally avoid them.

The absence of phrase repetition is only meaningful relative to the document job being performed.


What a Natural-Language Solution Must Explain

A natural-language decipherment should eventually make the repetition pattern unsurprising.

If duplicated tokens are reduplication, translate the grammatical effect consistently.

If exact phrase recurrence is suppressed by rich inflection, show the underlying repeated syntactic templates.

If spaces divide sub-word units, reconstruct the higher-level words and see whether phrase repetition reappears there.

If latent word classes carry syntax, class-level phrase recurrence should become stronger than exact-token recurrence.

The solution should not merely produce meaningful translations.

It should explain why the visible surface repeats at the levels it does.


What a Cipher Solution Must Explain

A cipher theory has a different route.

It can claim repeated plaintext phrases are deliberately diversified in ciphertext.

Then it should show:

  • how the same plaintext token can generate multiple visible forms;
  • why exact visible doubling still occurs sometimes;
  • why whole-word bigram predictability is weak;
  • why token-edge characters retain some cross-boundary dependency;
  • why line boundaries matter.

A historically plausible encryption procedure should reproduce this entire profile from readable plaintext.

That is a much stronger cipher test than simply finding one possible plaintext phrase.


What a Generation Solution Must Explain

A generation model can reproduce local repetition elegantly.

Its harder task is selectivity.

Why exact copy here?

Why one-edit mutation there?

Why not more repeated bigrams?

Why do repetition rates change across document roles?

Why are line boundaries respected?

A convincing generator should specify fixed probabilities or rules that reproduce the observed mix of exact and approximate repetition without tuning each passage individually.

Then it must still explain the larger document architecture.


The Strongest Repetition Model Must Explain Absence

This is the governing principle.

A model that can explain a doubled token is easy to construct.

A model that explains:

  • doubled tokens;
  • short local runs;
  • near-repetition;
  • line-boundedness;
  • few repeated exact bigrams;
  • few repeated exact trigrams;
  • weak whole-word predictability;
  • stronger edge-level dependency

is much more constrained.

The missing repetitions are therefore not empty space in the analysis.

They are negative evidence against mechanisms that would predict far more formulaic phrase reuse.


What Survives the Repetition Work

  • Exact token repetition exists in Voynich.
  • Local runs can contain several identical and near-identical forms.
  • Exact repeated multi-token phrases are comparatively scarce.
  • Whole-token bigram context provides surprisingly weak improvement in predicting the next exact token in at least major studied B material.
  • Token-edge characters still show some cross-boundary dependency.
  • Repeated sequences interact with line architecture.
  • Natural-language reduplication, records, cipher diversification, copying and local generation are all viable mechanism families for parts of the phenomenon.
  • No one mechanism explains the entire repetition profile yet.
  • The correct model must explain both observed repetition and suppressed repetition.

What Does Not Survive as Established Knowledge

  • Every doubled token is grammatical reduplication.
  • Every doubled token is a scribal error.
  • Repeated runs prove meaningless generation.
  • Repeated runs prove copying from the previous word.
  • Repeated tokens are proven quantities.
  • Three repeats form a universal dosage code.
  • Weak repeated bigrams prove there is no syntax.
  • Weak exact phrase recurrence rules out natural language.
  • Phrase suppression proves ciphertext.

The paradox survives.

The single grand explanation does not.


A Better Repetition Analysis

  1. Separate exact repeats from near-repeats.
  2. Count doubled tokens separately from repeated bigrams and trigrams.
  3. Control for base token frequency.
  4. Condition on line and paragraph position.
  5. Condition on Currier regime.
  6. Separate document roles.
  7. Track token-edge dependencies independently of whole-token bigrams.
  8. Test sensitivity to transcription.
  9. Compare language, record, cipher and generation models on the same repetition profile.
  10. Require the model to predict absences as well as repeats.

What Would Count as a Real Repetition Breakthrough?

Imagine a decipherment discovers that doubled tokens consistently encode one grammatical operation.

The rule predicts new doubled forms.

Underlying class-level syntax explains why exact full phrases seldom repeat even though grammatical templates do.

That would resolve the paradox linguistically.

Or imagine a historically plausible cipher maps repeated plaintext formulae into varying visible tokens while allowing exact duplication under one specific state.

The resulting ciphertext reproduces Voynich phrase scarcity, edge dependencies and line behaviour.

That would resolve it cryptographically.

Or a fixed local-generation process might reproduce the exact/near-repeat ratio, line confinement and bigram scarcity across unseen folios.

That would strengthen generation dramatically.

The breakthrough comes when one model predicts the entire repetition spectrum.


Primary School: Word Repeat and Phrase Repeat

Write ten short sentences in which the word “red” appears often.

Then ensure the exact phrase “red flower” appears only once.

Ask the child:

Is “red” repetitive? Is “red flower” repetitive?

The answers differ.

The child learns that repetition depends on the size of the unit being measured.


Lower Secondary: Same Vocabulary, Different Order

Give two groups the same twenty word cards.

Group A repeatedly uses the same phrases.

Group B shuffles the words under looser rules.

Both have the same word frequencies.

Their bigram repetition differs dramatically.

This teaches why unigram vocabulary statistics and phrase statistics are separate layers.


Upper Secondary: Compare Four Generators

Build four toy corpora.

  • ordinary sentences;
  • records with repeating fields;
  • ciphertext with homophones;
  • copy-and-mutate generated tokens.

Measure:

  • exact doubled words;
  • repeated bigrams;
  • repeated trigrams;
  • near-neighbour frequency.

The student can see that different mechanisms leave different repetition fingerprints.

This is exactly the comparison Voynich needs.


JC and Adult Readers: Think in Scale-Dependent Dependence

At a higher level, repetition shows that dependence changes with representation scale.

Character transitions are strong.

Token-internal families are strong.

Exact whole-token order is weaker.

Token-edge dependence persists.

Long-range topical distribution exists at another scale.

Voynich therefore cannot be understood from one Markov order or one unit type.

The manuscript’s information architecture may distribute dependence across several scales rather than concentrating it in ordinary word-to-word syntax.


A Parent and Teacher Guide

  1. Separate exact repeats from similar forms.
  2. Measure repetition at word, pair and phrase scale.
  3. Notice that frequent words do not guarantee frequent phrases.
  4. Ask whether repeats stay inside lines.
  5. Compare language, records, encoding and local generation.
  6. Do not turn every double into a grammatical or numeric meaning.
  7. Use missing expected repetitions as evidence too.

The transferable lesson is important:

what a system refuses to repeat can be as informative as what it repeats.


Reader Checklist: Before You Explain Voynich Repetition

  1. Is the repeat exact or only visually similar?
  2. Which transcription defines exact identity?
  3. Is it one doubled token, a bigram or a trigram?
  4. How common is the repeated token normally?
  5. Does the repeat stay inside one line?
  6. Does Currier regime affect repeat rates?
  7. Does document role affect them?
  8. Could the repeat be ordinary linguistic reduplication?
  9. Could it be a repeated record field?
  10. Could an encoding system diversify repeated plaintext?
  11. Could local copying or mutation generate the run?
  12. Does the theory also explain why exact phrases are scarce?

Frequently Asked Questions

Does the Voynich Manuscript repeat words?

Yes. Exact token doubling and short local runs occur, alongside the even more common phenomenon of near-related token families.

Does it repeat phrases?

Exact repeated two- and three-token sequences are surprisingly scarce relative to what one might expect from the recurrence of individual tokens and compared with several natural-language controls.

Does this prove generated text?

No. Local generation explains the pattern naturally, but language, record structure, encoding and copying can also contribute. The whole manuscript provides additional constraints.

Could doubled words be reduplication?

Possibly. Natural languages use reduplication, but a grammatical model would need to show systematic classes and effects rather than interpreting each repeat separately.

Could the repeats be scribal errors?

Some may be dittography or other copying errors. Systematic distributions would require a broader explanation.

Why is weak phrase repetition important?

Because ordinary meaningful prose often repeats common multi-word constructions. Voynich’s weak exact phrase recurrence suggests that its visible token order is organised differently or that the visible tokens are not ordinary plaintext words.

Is there still dependence between neighbouring tokens?

Yes. Character-level features at the end of one token can help predict the first one or two characters of the next, suggesting that the cross-boundary structure may operate below whole-token identity.

What is the strongest current conclusion?

Voynich repetition is scale-dependent: exact token repeats and local echoes exist, while exact longer phrase recurrence and whole-token next-word prediction are unusually weak. Any serious model must explain both sides of that pattern.


Related eduKateSG Reading


Research and Further Reading


The Final Idea

Voynich does not simply repeat too much.

It repeats at the wrong scale for an easy explanation.

A token echoes itself.

A nearby form almost echoes it.

A line can become locally repetitive.

And yet the manuscript does not flood itself with the same two-word and three-word phrases.

That tension is a clue.

Perhaps the real grammatical units are smaller than the tokens.

Perhaps an encoding process diversifies phrases.

Perhaps lines are generated from local material under rules that favour variation.

Perhaps the document is record-like rather than prose-like.

The final decipherment should not make us choose among these by taste.

It should make the strange repetition profile emerge naturally from one working system.

The manuscript repeats enough to show structure, but not in the way ordinary phrase memory would lead us to expect. That mismatch is one of its most useful constraints.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading