One of the most seductive Voynich ideas requires no hidden language at all.
Write a word.
Look nearby.
Copy something similar.
Change it.
Use the new form as the source for another form.
Continue long enough and the page begins to grow families.
daiin.
dain.
chedy.
chey.
Nearby words resemble one another.
Word families become networks.
Common forms become common because descendants keep citing them.
This is the intuition behind Torsten Timm and Andreas Schinner’s self-citation hypothesis.
It is one of the most serious meaningless-text mechanisms proposed for Voynich because it is not merely a story.
It has an executable generator.
It produces synthetic text.
Its output can be measured against the manuscript.
The important question is no longer “Can copy-and-modify make something that looks Voynich-like?” It can. The harder question is whether it reproduces the manuscript’s full behaviour strongly enough to be the historical mechanism.
Quick Read
- Timm and Schinner published a concrete self-citation generator in Cryptologia, with the article published online in 2019 and appearing in volume 44 in 2020.
- The core idea is recursive: previously written tokens become sources for later tokens.
- Modification rules include replacing glyphs with similar-shaped glyphs, adding or removing prefixes, and combining source words.
- The generated output reproduces several Voynich-like properties, including word-similarity networks and both Zipf laws in the authors’ reported tests.
- The model is historically attractive because it can in principle be executed manually without advanced machinery.
- The authors use these successes to support a meaningless-text or hoax interpretation.
- Critics have noted that early generator versions can overproduce unattested forms, misplace position-sensitive forms, and fail some detailed bigram or structural distributions.
- More importantly, reproducing one family of statistics does not establish that the historical manuscript was generated that way.
- A 2026 preprint by Rozanova and Temerev reports that a self-citation generator can reproduce several difficult Voynich properties—low entropy, multi-symbol unit scale, weak whole-token order and a failed simple-substitution attack—but still fails their edge-glyph coupling and open, hapax-rich vocabulary criteria.
- A separate 2026 directional study found that simple generator classes tested there failed a joint set of boundary and directional signatures.
- Therefore self-citation remains a strong control model and a serious mechanism hypothesis, but not a demonstrated origin story.
What “Self-Citation” Actually Means
The phrase can sound abstract.
The mechanism is concrete.
A token already present in the growing text becomes source material for a new token.
The new token is formed by rule-governed variation.
Then that new token joins the pool of material that can later be cited and modified.
The process is recursive.
Each generated token is both a result and a possible future source.
This creates historical memory automatically.
Early forms can become ancestors of large word families.
Nearby forms can resemble one another because the writer is drawing from recent or available material.
Rare mutations can either disappear or become productive if they are repeatedly reused.
The Modification Rules Matter
A sloppy summary says:
copy the previous word and change one letter.
That is too narrow.
Timm and Schinner describe a richer set of operations. Public documentation and their later discussion include rules such as:
- replace one or more glyphs with similar-shaped alternatives;
- add or remove a prefix;
- combine two source words to create a new word;
- reuse subgroups or repeated components under the model’s construction rules.
The order of glyphs is usually preserved rather than arbitrarily shuffled.
That matters because Voynich tokens are highly constrained internally.
A generator that changed characters freely would create huge numbers of legal-looking but unattested forms.
The self-citation hypothesis is interesting precisely because it tries to keep mutation inside a narrow historical path through a much larger theoretical word space.
Why Nearby Similar Words Matter
Timm’s earlier work emphasised a recurring observation: the more similar two Voynich tokens are, the more likely they are to occur near one another.
This can happen naturally under copying.
The eye remains in one part of the page.
A familiar form is reused.
A component is altered.
The new form appears close to its source.
Repeat the process and local neighbourhoods become family-rich.
This connects directly with the existing Word Families and Keywords Without Meanings evidence.
But local similarity has more than one possible cause.
- copy-and-modify generation;
- ordinary morphology;
- repeated technical vocabulary;
- cipher homophones;
- scribal spelling variation;
- local topic or document role.
Self-citation is therefore one causal model for a real observation.
Why Zipf Is Not Enough
The self-citation output can reproduce both of Zipf’s laws in the authors’ reported experiments.
That is an important demonstration.
It shows that a familiar language-like frequency law can arise from a mechanism that does not carry ordinary semantic prose.
This destroys a weak argument:
Voynich follows Zipf, therefore Voynich must be natural language.
The correct conclusion is smaller:
Zipf-like behaviour is compatible with multiple mechanisms.
That is why our Frequency Laws article treats Zipf as a property to explain rather than a language certificate.
Why Word-Length Fit Is Not Enough
The same lesson applies to Voynich’s narrow word-length distribution.
A constrained generator can produce a narrow token-length envelope if its allowed operations preserve a small family of structural templates.
Therefore matching the histogram is useful.
It is not mechanism identity.
The Word-Length Problem exists precisely because several very different systems can create similar length profiles.
Why the Network Result Is More Interesting
Timm and Schinner represent word types as a network connected by similarity relations.
Their generated text produces a large connected network with path-length and coverage properties similar to the manuscript under their chosen definitions.
This matters because the Voynich vocabulary is not merely a bag of unrelated forms.
It has neighbourhood structure.
Many tokens sit close to other tokens under edit-like or component-like similarity.
A recursive generator naturally creates such genealogies.
Again, the causal ambiguity remains.
Languages with productive morphology also create families.
Ciphers with verbose alternatives can create families.
Abbreviation systems can create families.
The network is real.
The historical cause is still open.
The Strongest Version of the Self-Citation Claim
The strongest claim is not simply that the model imitates statistics.
It is that the manuscript itself was generated by this kind of process and therefore lacks ordinary semantic plaintext.
That is a much larger claim.
To earn it, the model must explain manuscript-wide structure beyond the properties used to motivate the generator.
- Currier A/B or continuous drift;
- scribe/hand effects;
- line-position effects;
- paragraph-initial Grove behaviour;
- labels versus running text;
- circular and radial loci;
- edge and boundary coupling;
- rare-glyph distributions;
- open vocabulary growth;
- section-local vocabulary;
- visual/text relationships.
A surface generator becomes a historical theory only when it survives those independent dimensions.
The Early Critique: Overgeneration
One criticism raised in specialist discussion is that a flexible copy-and-modify process can generate too many forms.
If a source token can mutate in many legal ways, why does the manuscript occupy such specific narrow paths through the possible vocabulary?
Why are many edit-distance-near forms absent?
Why do certain bigrams remain strongly suppressed?
Why are some forms restricted to line starts or other positions?
Timm responds by making the modification rules more structured than arbitrary mutation. Similar-shaped glyph substitutions, prefix operations and source-word combinations restrict the path.
This is exactly where falsifiability lives.
The more rules are specified in advance, the more strongly the generator can be tested.
The Position Problem
Voynich tokens do not appear equally everywhere.
Some forms prefer line beginnings.
Some endings prefer line ends.
Paragraph openings favour gallows-rich forms.
A generator that selects a source and mutates it without accounting for document position can reproduce the vocabulary while misplacing the vocabulary.
This is a deeper failure than a wrong frequency.
It means the mechanism does not yet know where it is on the page.
The manuscript does.
The 2026 Boundary Challenge
The newest important stress test comes from a 2026 preprint by Liudmila Rozanova and Alexander Temerev.
Their study explicitly compares Voynich against prose, cipher and pseudo-text controls and tests assumptions about glyphs, tokens and spaces.
Crucially, they report that a self-citation generator reproduces several difficult Voynich properties:
- low character entropy;
- a recurrent multi-symbol unit scale;
- weak predictive order at the exact whole-token level;
- failure of a calibrated simple-substitution attack.
That is a significant success for self-citation as a control.
Then come the failures.
Under their analysis, self-citation does not reproduce the manuscript’s reported coupling between glyphs at token edges.
It also does not reproduce the manuscript’s open, singleton-rich vocabulary to the same degree.
The authors report roughly 70% singleton types in Voynich against substantially lower proportions in their self-citation and cipher controls.
The precise numbers belong to a preprint and should be replicated.
The methodological lesson is already strong:
a generator can pass old Voynich tests and still fail new ones that examine a different layer.
The 2026 Directional Challenge
Parisel’s 2026 directional preprint creates another higher bar.
His analysis reports a directional split between token-internal character structure and cross-token boundary dependence.
The tested simple generator families did not reproduce the complete joint signature across their tested parameter spaces.
This does not test every implementation of Timm–Schinner self-citation directly.
It does tell all local-copy models what the next exam looks like.
It is no longer enough to create similar nearby tokens.
The process must also create the correct directionality inside and between them.
See The Directional Dissociation Problem.
Long-Range Memory Is Both a Strength and a Problem
A recursive generator naturally creates memory.
Current tokens descend from earlier tokens.
Shuffling tokens can therefore destroy long-range statistical structure.
Recent community analyses have compared this behaviour with natural-language controls, the Naibbe cipher and Timm-generated text.
Some aspects of Voynich long-range mutual information resemble the self-citation output more than ordinary shuffled prose.
Other aspects remain different.
This is exactly why the next article in this batch owns long-range memory separately.
A mechanism may explain local families and still get manuscript-scale persistence wrong.
The Historical Plausibility Question
Self-citation has one major historical advantage.
It needs no computer.
A scribe can look at nearby material.
Copy.
Vary.
Repeat.
The fact that a mechanism is medievally executable matters.
But historical possibility has the same asymmetry we meet everywhere in Voynich research.
“A fifteenth-century scribe could do this” is weaker than “this fifteenth-century manuscript was made this way.”
Meaningless Text Is Not the Only Way to Use Copy-and-Modify
There is another important logical distinction.
Even if copy-and-modify behaviour were historically real, it would not automatically prove semantic emptiness.
Meaningful systems also reuse and vary forms.
- morphology modifies stems;
- abbreviations create families;
- cipher homophones create alternate visible forms;
- scribes copy formulaic phrases and alter local details;
- technical registers repeat templates.
The same observable similarity network can therefore arise from different semantic states.
A self-citation-like operation and a meaningless-text interpretation are separable claims.
The Better Use of Self-Citation: Keep It as a Control
This is where the model becomes most valuable regardless of whether the historical hypothesis eventually wins.
A known generator gives us a positive control.
We know the ground truth:
- the output is generated;
- the mechanism is recursive;
- nearby similarity is causal;
- there is no hidden natural-language plaintext in the generated sample.
Now any proposed Voynich statistic can be tested against that known mechanism.
If a statistic labels self-citation output “natural language”, the statistic is weak.
If a statistic cleanly separates real Voynich from self-citation under fair preprocessing, it becomes a candidate discriminator.
This is the deepest value of a serious failed model.
A mechanism does not need to be historically correct to become scientifically useful.
What Would Make Self-Citation Much Stronger?
- Independent implementations reproduce the authors’ core results.
- The mechanism reproduces line-start and paragraph-start effects without page-specific hand tuning.
- It reproduces Currier A/B or continuous drift under one coherent generative history.
- It matches edge-glyph coupling and graded-space behaviour.
- It matches singleton-rich vocabulary growth.
- It reproduces directional dissociation.
- It preserves visual/document-role differences between labels and prose.
- Its parameters are fixed on one subset and succeed on held-out folios.
What Would Weaken It?
- The strongest similarities disappear under better glyph segmentation.
- More constrained natural-language or cipher controls reproduce the same networks equally well.
- The generator requires growing lists of local exceptions to reproduce independent Voynich features.
- Held-out pages systematically fail after parameters are frozen.
- Observed directional or boundary signatures are incompatible with the recursive modification process.
Primary School: Copy, Change, Repeat
Write the made-up word:
lom
Now make:
- loma;
- lomar;
- romar;
- rom;
- trom.
After twenty rounds, the class has a vocabulary full of related forms even though none has been assigned meaning.
That demonstrates why word families alone do not prove language.
Secondary School: Build the Generator and Then Try to Catch It
Create two pages of copy-and-modify pseudo-text.
Measure:
- word-frequency rank;
- word length;
- nearby similarity;
- number of unique types.
Then design a test that distinguishes the synthetic text from real prose.
Students learn the difference between reproducing a statistic and reproducing a mechanism.
JC and Adult Readers: Treat It as a Generative Model, Not an Argument
Generate synthetic corpora under frozen rules.
Compare them with Voynich on metrics not used to design the generator.
- edge coupling;
- directional dissociation;
- hapax growth;
- Currier transfer;
- line-position behaviour;
- long-range MI;
- label/prose differences.
The method should win because it predicts unseen structure, not because it was built to imitate visible structure.
Reader Checklist: Before You Say Self-Citation Explains Voynich
- Which exact self-citation rules are being used?
- How is a source token selected?
- What modifications are permitted?
- Which statistics were used to design the generator?
- Which tests are genuinely held out?
- Does the model reproduce position-specific forms?
- Does it reproduce Currier variation?
- Does it reproduce edge and space behaviour?
- Does it reproduce open vocabulary growth?
- Does it reproduce directional dissociation?
- Does the result imply generation only, or generation plus meaninglessness?
- Could a meaningful cipher or abbreviation system create similar copy-like families?
Frequently Asked Questions
What is the Voynich self-citation hypothesis?
It is Timm and Schinner’s proposal that a scribe generated Voynich-like text recursively by copying previously written tokens and modifying them under constrained rules, with generated tokens becoming sources for later ones.
Does it reproduce Voynich statistics?
Yes, several important ones. The published model reproduces strong word-similarity-network behaviour and both Zipf laws in the authors’ reported tests. Newer work also finds that self-citation controls can reproduce some difficult low-level properties.
Does that prove Voynich is meaningless?
No. It proves that some famous Voynich properties can arise without ordinary semantic prose. The historical manuscript still has additional structure the mechanism must explain.
Is the generator historically possible?
The core copy-and-modify procedure is simple enough to be carried out manually. Historical executability is a strength, but it is not evidence that the actual Voynich scribes used the procedure.
Why keep testing it?
Because it is one of the most useful known-mechanism controls in Voynich research. Any statistic claimed to distinguish meaningful structured text from generated pseudo-text should be tested against it.
Research Foundations
- Torsten Timm & Andreas Schinner — A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1).
- Timm — Self-Citation Text Generator: public additional materials and code.
- Rozanova & Temerev — A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space (2026 preprint).
- Parisel — Evidence of Layered Positional and Directional Constraints (2026 preprint).
The Final Idea
Self-citation deserves respect because it made Voynich research harder in the right way.
It showed that a meaningless generator can grow families, networks and frequency laws that look impressively language-like.
That destroyed several easy arguments.
Now newer measurements are doing the same thing back to self-citation.
They are asking it to explain edges.
Direction.
Open vocabulary.
Long-range memory.
A good Voynich theory should not merely survive the tests available when it was invented. It should survive the tests invented because it existed.