A book can talk about the world.
It can also talk about itself.
See page 24.
Use the same ingredient as above.
Refer to figure B.
Repeat the preparation listed under another entry.
Modern books make such references explicit.
Medieval manuscripts can do it more quietly.
A repeated rubric.
A sign.
A marginal symbol.
A repeated label.
A known abbreviation.
A copied identifier.
Voynich contains hundreds of labels and thousands of recurring forms.
Some labels also appear, or nearly appear, in running text.
Some rare forms recur across distant visual sections.
Some modern computational projects now claim cross-section label networks linking visually matching objects.
This is potentially enormous.
If one token really identifies the same internal entity on two distant pages, we gain semantic structure without needing to know the entity’s modern name.
But repetition is cheap.
Common words repeat.
Morphology repeats.
Cipher equivalents repeat.
Self-citation generators repeat and modify their own material.
Scribes reuse familiar shapes.
A document can therefore contain recurring tokens without any one of them meaning “go look somewhere else”.
The Cross-Reference Problem asks what extra evidence turns recurrence into intentional reference.
Quick Read
Direct answer: the Voynich Manuscript contains enough repeated labels, rare tokens and cross-section recurrences to make internal reference a serious research question, but no manuscript-wide cross-reference system has been accepted. A true cross-reference should do more than repeat: it should preserve identity or relation across distance, be distinguishable from ordinary vocabulary and copy-and-modify generation, align with independent visual or document evidence, and predict unseen links under a frozen rule.
- Jorge Stolfi mapped possible exact and near references from Voynich labels into running prose as early as the 1990s.
- Stolfi explicitly warned that many near-matches may be bogus.
- Exact label recurrence is more informative than an arbitrary near-match but still does not prove intentional reference.
- A common token can recur because it is common language.
- A rare token recurring in matched visual contexts is much more interesting.
- A repeated label can represent the same object, same class, same function or merely a reused writing template.
- Modern computational work now claims cross-section label networks linking some repeated tokens to visually similar objects.
- Those claims remain ongoing research rather than accepted decipherment.
- The Self-Citation Problem owns copying and modifying nearby words as a production mechanism; this article owns documentary reference across distance.
- A cross-reference does not need to contain a page number.
- A stable identifier can function as an internal link.
- Cross-reference can ground internal identity before real-world meaning is known.
- Cipher homophones can hide references by allowing the same underlying identifier several visible forms.
- Currier A/B can change surface forms and therefore complicate exact-match networks.
- A real reference system should improve retrieval and predict where related entries occur.
- No accepted internal index or cross-reference grammar has yet been recovered.
Reference Is More Than Recurrence
Suppose the same Voynich token appears on f20r and f90v.
We know one thing.
The same visible form occurs twice.
We do not yet know why.
- same word;
- same identifier;
- same grammatical marker;
- same cipher form;
- same template;
- copying coincidence.
Reference is a relational claim.
Occurrence B points back to, identifies, invokes or reuses the same external/internal entity as occurrence A.
That relationship requires evidence beyond identity of spelling.
The Strongest Cross-Reference Is an Identifier
Imagine one rare label appears beside a distinctive root on one page.
The same rare label appears beside a highly similar root elsewhere.
Then a related visible component appears in running text near a pharmaceutical entry.
This creates a three-way relation:
- textual identity;
- visual identity;
- document-role recurrence.
If these relationships repeat across many independent cases, the token begins to function like an identifier.
We may still not know the object’s modern botanical name.
But we have grounded internal identity.
Stolfi’s Label Reference Maps Were an Early Serious Attempt
Jorge Stolfi searched hundreds of transcribed Voynich labels for exact and near occurrences in the manuscript’s prose.
The resulting maps remain valuable because they frame the problem operationally.
Where is a label defined?
Where else does the same form occur?
How dense are possible references by page?
Stolfi also included the essential warning:
many near-matches are probably bogus.
That warning remains central almost thirty years later.
Why Near-Matches Are Dangerous
Voynich words live in dense similarity neighbourhoods.
Change one glyph and another valid-looking word often appears.
This means a fuzzy search can always find relatives.
Suppose label otaly has dozens of near-neighbours.
If every neighbour counts as a reference, the network becomes dense automatically.
A real reference map therefore needs a calibrated similarity threshold.
How many links appear in shuffled or matched-control text under the same threshold?
That null network is essential.
Exact Matches Are Better, but Still Not Enough
Exact identity removes one degree of freedom.
Good.
But exact words in natural language recur too.
So do common cipher groups.
The stronger cross-reference candidates are:
- rare exact labels;
- occurrences in role-compatible contexts;
- links across distant pages;
- links accompanied by independent visual correspondence.
Rare Tokens Carry More Reference Information
If a token appears a thousand times, two matching occurrences tell us little.
If a token appears exactly twice, on distant pages, beside visually similar rare objects, the coincidence is much harder to dismiss.
This is why hapax and rare-token statistics matter to cross-reference analysis.
Rare words are potential identifiers.
They are also potential transcription errors and local mutations.
Both possibilities must remain.
The Modern Cross-Reference Claim
An ongoing 2026 public computational project by G. Taghon reports a label cross-reference network linking specific labeled objects across several applied Voynich sections.
The project describes repeated labels associated with visually matching illustrations and presents a preliminary lexicon with graded confidence.
This is exactly the kind of claim worth testing.
It is not yet a settled result.
The strongest contribution is methodological:
if text and visual identity independently converge across distant pages, internal reference becomes testable before full translation.
Cross-Reference Is Not Self-Citation
The words sound similar.
The mechanisms are different.
Self-citation
A writer looks at existing Voynich words and copies or modifies them to generate new surface forms.
The relationship is production genealogy.
Cross-reference
A form intentionally refers to the same entity, class, record or operation elsewhere.
The relationship is documentary meaning or retrieval.
A self-citation generator can create lookalike “reference” networks with no semantic identity.
Therefore any cross-reference metric must be tested against self-citation output.
Cross-Reference Is Not Topic Recurrence
Two plant pages can share vocabulary because both discuss plants.
That does not mean one refers to the other.
Topic creates broad statistical similarity.
Reference creates a narrower link.
The same named ingredient.
The same procedure.
The same entry code.
A good model should distinguish these resolutions.
Cross-Reference Is Not Morphological Family
Two related forms may share a root and differ by suffix.
That is a morphological relationship.
It does not imply they point to the same entity.
One may be singular.
Another plural.
One noun.
Another adjective.
Reference analysis should either work at the underlying lemma level under a stable morphological model or restrict itself to exact identifiers.
Homophony Can Hide True References
If the manuscript uses homophonic ciphering, the same plaintext name may have several visible forms.
Exact-match networks will miss some real references.
This is why Averyanov’s zodiac-sigla work tests collapsing labels by proposed homophone classes before concluding that ring vocabularies remain distinct.
The lesson generalises.
Reference analysis must either:
- stay at visible exact form and accept low recall;
- or use independently recovered equivalence classes without inventing them from the desired cross-reference.
Currier A/B Can Hide True References Too
If A and B represent different orthographies, cipher states or shorthand conventions, one underlying identifier may change surface form.
A cross-reference system should therefore test whether A↔B transformations increase meaningful links under a rule learned elsewhere.
If the transformation is chosen because it makes one desired link appear, the evidence is circular.
Labels Are More Likely to Be Reference Handles Than Prose Words
Labels are compact.
Anchored.
Often unique.
Those are properties of identifiers.
A practitioner does not need a full sentence to cross-reference an ingredient.
A short code is enough.
The Audience and Notation problems therefore make cross-reference especially plausible in a specialist reference book.
The Genre Problem Predicts Whether Cross-References Should Exist
A linear literary text may need few internal links.
A reference manual benefits enormously from them.
A medical compendium may reuse ingredients across recipes.
A catalogue may point from whole plants to detached components.
A zodiac index may connect names to timing entries.
Therefore the existence and density of real cross-references can become a genre discriminator.
Whole Plant to Fragment Is the Obvious Internal Corridor
Voynich has large plant-dominated pages and later pages containing detached plant-like components.
This creates a natural cross-reference hypothesis.
Does a fragment page point back to the corresponding whole plant?
This is stronger than a speculative whole-book pipeline because the relation is local and testable.
- match visual component;
- match label or identifier;
- test whether matches exceed chance;
- check unseen pages.
The corridor may survive even though the larger plant→body→zodiac→recipe story failed.
Visual Matching Must Be Independent
Suppose a researcher finds token X on two pages.
They then search the illustrations until two vaguely similar roots appear.
That is weak.
The visual match was selected after the textual match.
A stronger design uses blind image matching.
- Freeze textual candidate links.
- Hide the labels from visual assessors or algorithms.
- Rank image similarity independently.
- Ask whether linked pages are unusually high in that ranking.
Now one modality validates the other.
A Cross-Reference Should Be Directional or Role-Specific
Not every internal link needs an arrow glyph.
But role can create direction.
A whole-plant page may define an entity.
A pharmaceutical fragment page may reuse its identifier.
A recipe may refer to it again.
If one document role consistently introduces a label and another reuses it, reference becomes stronger.
The introduction/reuse asymmetry is a useful test.
Current Folio Order Should Not Define Reference
Voynich has been rebound.
Leaves are missing.
Modern adjacency is not always original adjacency.
A true internal identifier should work regardless of current page numbering.
In fact, cross-reference networks may help reconstruct original organisation if their direction aligns with bifolia and quires.
But this inference must not be circular: do not reorder pages to maximise links and then cite the increased links as proof of the order.
Cross-Reference Can Help Find Missing Leaves
Suppose several entries refer to an identifier whose defining page is absent.
This could imply a missing source entry.
That is a powerful prediction.
A mature network could therefore identify “dangling references”.
But the model needs established introduction/reuse roles before absence becomes evidence.
Cross-Reference Can Ground Semantics Without Translation
This is the deepest value.
Suppose token X consistently tracks the same internal plant entity.
We do not know whether X is the plant’s name.
It could be an index code.
But we know something semantic:
X preserves identity across contexts.
This is internal reference.
It can be discovered before real-world naming.
Reference Networks Should Be Sparse Enough to Be Useful
If every page links to every other page, the network explains nothing.
A useful reference system is selective.
It creates communities.
Hubs.
Definitions.
Reuse points.
A network derived from permissive fuzzy matching will often be too dense.
Sparsity under strong evidence is a feature, not a failure.
The Correct Null Model Is Hard
Randomly shuffling Voynich words may be too easy a control because it destroys the manuscript’s real local families.
A stronger null preserves:
- token frequency;
- word-family density;
- Currier state;
- document role;
- section vocabulary;
- label length.
Then ask whether candidate reference links still exceed expectation.
Self-citation output is another important matched control because it generates local genealogies without semantic identity.
A Real Cross-Reference System Should Improve Retrieval
Imagine a trained reader sees a label on a pharmaceutical page.
If that label is an identifier, they should be able to locate the corresponding whole plant or another related record.
That is the user story.
A historical cross-reference system exists because it saves work.
Genre and audience therefore become part of the evidence.
A reference-heavy architecture fits a specialist manual better than a continuous literary text.
What Survives the Cross-Reference Work
- Voynich labels and rare forms recur across distant parts of the manuscript.
- Stolfi’s early label maps established that exact and near label-to-prose links can be systematically searched.
- Near-match networks have high false-positive risk because Voynich word families are dense.
- Rare exact matches in independently matching visual contexts are stronger reference candidates.
- Modern computational projects now claim cross-section label networks, but these are not accepted decipherments.
- Cross-reference is distinct from self-citation, topic recurrence and morphology.
- Homophony and Currier state may hide true references behind alternate surface forms.
- Whole-plant ↔ fragment pages form a particularly strong corridor for testing internal identity.
- Cross-reference can ground internal identity before real-world naming.
- No accepted manuscript-wide cross-reference grammar or index has been recovered.
What Does Not Survive as Established Knowledge
- Every repeated Voynich label is a cross-reference.
- Every label occurrence in prose refers back to the labelled picture.
- Near-matching tokens are references.
- The manuscript has a recovered internal index.
- One shared token proves two drawings depict the same real-world object.
- The Taghon cross-reference network is accepted semantic decoding.
- Current folio order defines reference direction.
- Cross-section token recurrence proves the failed plant→body→zodiac→recipe pipeline.
A Better Cross-Reference Analysis
- Begin with exact rare label matches.
- Separate labels, prose and record-like text.
- Control token frequency and word-family density.
- Freeze textual links before image comparison.
- Use blind visual similarity or independent image classifiers.
- Test self-citation and shuffled-family controls.
- Learn any homophone or Currier transformations independently.
- Look for introduction versus reuse asymmetry by document role.
- Validate on unseen pages and physical units.
- Require the network to improve retrieval or predict missing/related entries.
What Would Count as a Real Cross-Reference Breakthrough?
Imagine a set of rare labels is frozen from one part of the manuscript.
Without looking at images, the model predicts where those labels should recur.
Independent blind image analysis then finds that linked pages contain matching rare visual components far above matched-control expectation.
One role consistently introduces identifiers and another reuses them.
The same rule works across Currier states under a previously learned transformation.
Several predicted links land on pages not used to build the network.
One predicted target is missing at a known lacuna, creating a falsifiable reconstruction.
That would be an internal reference system.
The manuscript begins to point to itself when a token does more than recur—when it carries the same identity across distance strongly enough to find the right place before we look.
Primary School: Same Sticker, Same Thing?
Put the same invented sticker on two classroom objects.
Does the sticker mean both objects are identical?
Maybe it marks “science equipment”.
Maybe it marks “belongs to Group A”.
Students learn that repeated labels establish a relation before they establish the exact relation type.
Secondary School: Build a Reference Network
Create an invented manual with repeated identifiers across diagrams and recipes.
Then mix in ordinary repeated words.
Ask students to recover only the true identifier links.
They quickly discover why rarity, role and independent visual agreement matter.
JC and Adult Readers: Cross-Reference as Graph Inference
At a higher level, pages and labelled entities form nodes.
Candidate textual identities form edges.
The task is to infer a sparse latent reference graph from a noisy graph of lexical similarity.
Independent visual evidence acts as a second edge channel.
A true reference network should show cross-modal edge enrichment beyond matched null graphs while retaining predictive value on unseen nodes.
Reader Checklist: Before You Call a Voynich Match a Cross-Reference
- Is the token an exact match or fuzzy match?
- How frequent is it?
- Is it a label or running-text token?
- Could morphology explain the similarity?
- Could self-citation explain it?
- Could homophony hide alternate forms?
- Are the visual objects matched independently?
- Does the same relation repeat across several cases?
- Is one occurrence an introduction and another a reuse?
- Does the network survive Currier and hand controls?
- Does it predict unseen links?
- What retrieval job would the reference perform for the intended reader?
Frequently Asked Questions
Does Voynich contain internal cross-references?
Possible label and token reference networks have been identified, but no manuscript-wide system is accepted as demonstrated.
Did Stolfi find cross-references?
He systematically mapped exact and near occurrences of labels in running text, creating an important empirical foundation while explicitly warning that many near-matches may be bogus.
Is cross-reference the same as Timm’s self-citation?
No. Self-citation is a proposed text-generation process based on copying and modifying existing words. Cross-reference is an intentional documentary relation that preserves identity or points to related information.
Why are labels important?
Labels are short, anchored and often rare, making them more plausible identifier candidates than arbitrary common prose tokens.
Can cross-references help without translation?
Yes. If one identifier reliably tracks the same internal entity or class across pages, internal identity can be recovered before the real-world name is known.
What is the strongest current conclusion?
Voynich contains enough label recurrence and internal visual reuse to justify systematic cross-reference testing, but deliberate reference has not yet been demonstrated at manuscript scale.
Related eduKateSG Reading
- The Semantic-Grounding Problem
- The Self-Citation Problem
- Labels Versus Running Text
- The Herbal Pages
- The Vessels and Plant Fragments
Research and Further Reading
- Jorge Stolfi — Label Reference Maps
- René Zandbergen — Voynich Labels
- G. Taghon — Ongoing Voynich natural-language and cross-reference analysis (2026)
- Rozanova & Temerev — Representation and Control Problems in Voynich Analysis (2026 preprint)
The Final Idea
The great promise of cross-reference is that the manuscript may be able to define itself before we can translate it.
A root appears here.
A fragment appears there.
The same rare identifier follows both.
A later entry reuses it.
Now we have a relationship that does not depend on guessing whether the object is mandrake, violet or something else.
The book knows what it considers the same.
Our job is to prove that the repeated sign is carrying that identity rather than merely repeating because Voynich repeats.
If a genuine cross-reference network exists, it may become the first internal map from unreadable signs to stable entities—a map drawn by the manuscript itself.