Part I of X — The Object Before the Mystery
There is a particular kind of mystery that survives by being vague. The less we know, the more freely we can imagine. The Voynich Manuscript is not that kind of mystery.
It is much stranger than that.
We know a great deal about the object. We can measure it. We can examine its parchment. We can analyse samples of its inks and pigments. We can see where leaves have disappeared. We can reconstruct the logic of gatherings and foldouts. We can follow parts of its journey through seventeenth-century Europe, through the rare-book trade, and eventually into Yale University. We can count its pages, inspect its writing, compare one region of text with another, separate later annotations from the main script, and observe that some kinds of text behave differently from other kinds.
And after doing all of that, we still cannot read it.
That is the real intellectual beauty of Voynich. It is not a blank wall. It is a landscape in which many roads are visible and none has yet been shown, convincingly and reproducibly, to reach the destination.
This master article is eduKateSG’s attempt to describe that landscape properly. It does not begin by asking for a solution. It begins with a more disciplined question:
What can the surviving object itself tell us before we ask it to tell us what it means?
That question matters because Voynich attracts explanations faster than almost any historical object attracts evidence. A strange plant becomes a species identification. A circular diagram becomes a cosmology. A cluster of women and pools becomes a medical procedure. A repeated sign becomes a letter. A statistical pattern becomes proof of language. A resemblance to another manuscript becomes a provenance claim. A place that fits one clue becomes the birthplace of the whole book.
Sometimes those ideas are useful. Sometimes they are brilliant. Sometimes they are wrong. Usually, at the beginning, we do not yet know which.
Quick Read
The Voynich Manuscript is a surviving European parchment codex, now Yale’s Beinecke MS 408, whose parchment has been radiocarbon-dated to the early fifteenth century and whose unidentified writing system remains unread despite more than a century of intensive study.
- The surviving codex contains about 102 folios, with evidence that additional leaves once existed.
- Its standard leaves are roughly 23 by 16 centimetres, but the manuscript also contains unusually elaborate foldouts.
- Protein analysis identified the parchment as calfskin.
- Radiocarbon testing of four samples places the parchment broadly in the period 1404–1438 at the commonly reported 95% confidence range.
- Scientific examination found writing and drawing materials compatible with historical manuscript production rather than a simple twentieth-century fabrication.
- The present binding, folio order and surviving leaf sequence should not automatically be treated as the author’s original conceptual order.
- Fourteen folio numbers are absent from the surviving sequence; some leaves appear to have been cut out while other losses involve bifolia or entire gatherings.
- The earliest secure documentary discussion comes from seventeenth-century Central Europe, although some reported ownership claims reach back to the court of Rudolf II.
- Wilfrid Voynich acquired the manuscript in 1912. It eventually reached Yale, where it became Beinecke MS 408.
- Nothing about the physical dating identifies the language, author, purpose or exact place of production.
The single most important habit for reading this article is therefore simple: do not ask one piece of evidence to do the work of five.
How This Ten-Part Master Works
This is one article, built as ten large movements. Each part has its own job, but all ten serve one purpose: to preserve what the Voynich evidence has actually taught us so new readers do not have to restart the same circular arguments from zero.
- The Object Before the Mystery — parchment, dating, inks, pigments, quires, missing leaves, binding, provenance and the difference between an old object and a solved object.
- The World Inside the Book — plants, zodiac figures, stars, circles, the Rosettes foldout, human figures, pools, tubes, vessels, plant fragments and starred entries.
- What Voynichese Actually Does — transcription, EVA, glyphs, spaces, gallows, q-series forms, labels, lines, paragraphs and positional behaviour.
- The Text Is Not One Uniform Thing — Currier A and B, scribal variation, page families, local regimes and the hidden variables that can make one manuscript behave like several.
- Words That Behave Strangely — near-neighbour word families, local repetition, vocabulary geography, predictability, entropy and the difference between pattern and meaning.
- What the Pictures and Text Do Together — layout, labels, image-text relationships, recurrence across scales and what cross-modal structure can constrain without translating.
- The Manuscript World Around Voynich — Carrara, Padua, Masson 116, Sloane 4016, Egerton 747, medieval medicine, herbals, astrology, baths and the disciplined use of comparators.
- What We Tried, What Broke, and What We Should Stop Repeating — reader-safe lessons from attractive explanations that failed, overreached or refused to generalise.
- What Actually Survived — the strongest physical, scribal, textual, visual, geometrical and distributional findings that remain standing after weaker stories are removed.
- Where Voynich Research Should Go Next — the evidentiary burden for a genuine solution, a practical anti-loop checklist, and the frontier that still deserves work.
The growing eduKateSG Voynich library is not being replaced. Those specialist articles are the rooms around this central hall. This page is the master doorway, the synthesis, and the place where conclusions are frozen so readers do not have to rediscover the same dead ends.
Our Evidence Ladder
Voynich becomes easier to think about when claims are sorted by strength.
Known
Directly observable or supported by strong physical or documentary evidence. The folio exists. A leaf is missing. A material was detected in a tested sample. A letter survives in an archive. A radiocarbon measurement produced a probability distribution.
Strongly supported
The evidence is convergent, but the conclusion still contains interpretation. For example, a group of pages may show a coherent production relationship even if we cannot yet name the subject they encode.
Plausible
The proposal fits what we know and has a historically reasonable mechanism, but alternatives remain alive.
Possible
Nothing decisive rules it out, but the supporting evidence is weak, incomplete or highly non-unique.
Unsupported
An idea may be imaginative or even historically conceivable, but the manuscript has not yet supplied enough independent evidence for it.
Contradicted
The proposal conflicts with established evidence or requires so many exceptions that it stops functioning as an explanation.
This ladder is not a device for killing curiosity. It is how curiosity survives contact with reality.
Before It Was a Mystery, It Was a Book
Modern readers meet Voynich as an image on a screen. That is already a distortion.
The original is a codex: an engineered physical object made from animal skin, cut into sheets, folded, nested, written upon, illustrated, painted, gathered, sewn, rebound, numbered, handled, transported, owned, examined, damaged and preserved.
That sounds obvious until we notice how many theories quietly treat the manuscript as if it were a perfectly ordered modern PDF whose page sequence, section divisions and visual categories were guaranteed by its creator.
They are not.
The distinction matters because a codex has several histories at once. There is the history of the parchment. The history of the writing. The history of the illustrations. The history of the paint. The history of assembly. The history of binding. The history of numbering. The history of ownership. The history of loss. The history of modern study.
Those histories overlap, but they are not automatically identical.
What Exactly Is the Voynich Manuscript?
Today the object is held by the Beinecke Rare Book and Manuscript Library at Yale University and catalogued as Beinecke MS 408. The name “Voynich Manuscript” comes not from its maker but from Wilfrid M. Voynich, the Polish-born antiquarian bookseller who acquired it in 1912 and made it famous in the modern rare-book world.
The manuscript is small enough to feel personal rather than monumental. McCrone Associates measured the codex at approximately 23.5 cm high, 16.2 cm wide and about 5 cm deep during its 2009 materials examination. René Zandbergen’s codicological description gives a similar standard folio scale of roughly 23 by 16 cm. In other words, this is not a giant ceremonial atlas. It is a book that could be held, carried and worked with.
But “small book” does not mean simple book.
Some leaves unfold beyond the normal page dimensions. Some are multiple-panel foldouts. The famous large sheet associated with folios 85 and 86 opens into an unusually complex surface. Across the codex, drawings and text interact in ways that repeatedly force the reader to ask whether the page is being used as a page, a diagram, a table, a map-like surface, a labelled object field, a working sheet, or some combination of those things.
Already, before a single character has been “read”, the object is telling us something: its makers were willing to alter ordinary page geometry when the information they wanted to place on the surface demanded it.
Parchment Is Evidence
The leaves are parchment. More specifically, later protein-based testing identified the animal source as calfskin, so “vellum” is also technically appropriate in the stricter sense of that term.
This matters for several reasons.
First, parchment is manufactured material. A manuscript maker did not walk into a forest and find a blank Voynich page growing on a tree. Animals had to be obtained. Skins had to be processed, stretched, scraped, dried and cut. A codex of this size therefore represents material cost, labour and planning before the first line of strange script appears.
Second, parchment remembers production. Thickness changes. Holes and flaws survive. Hair-side and flesh-side characteristics may remain visible. Fold patterns, stubs, sewing structures and neighbouring leaves can reveal how sheets were once related.
Third, parchment can be dated by radiocarbon analysis. That does not solve Voynich, but it dramatically shrinks one part of the possibility space.
1404–1438: What the Radiocarbon Date Really Says
In 2009, four parchment samples from different parts of the manuscript were tested at the University of Arizona. The results were mutually consistent. The dating is commonly summarised as a 1404–1438 range for the parchment at approximately 95% confidence, with slightly different upper limits appearing in later recalculations depending on calibration treatment.
It is one of the most important pieces of evidence in the entire Voynich story.
It is also one of the easiest to misuse.
Radiocarbon dates the animal material, not the sentence
The measurement concerns when the biological material stopped exchanging carbon with the environment. It does not place a tiny historical clock next to the scribe’s pen.
In principle, parchment could be stored before use. Therefore the safest statement is not “the manuscript was written in 1421” or any similarly precise claim. The safe statement is that the sampled parchment belongs to the early fifteenth-century window indicated by the measurements, and any proposed production history must accommodate that fact.
The date destroys some theories more effectively than it proves others
Wilfrid Voynich promoted a possible connection with Roger Bacon, the thirteenth-century English Franciscan scholar. Once the parchment itself is placed in the fifteenth century, an original manuscript written by Bacon on these leaves becomes chronologically impossible. Bacon died long before the calves from which the parchment was made were born.
Notice the asymmetry. The date can rule out Roger Bacon as the physical author of this codex much more strongly than it can identify the real author.
This is a recurring pattern in Voynich research. Evidence often removes possibilities more efficiently than it names the surviving one.
The four samples matter
The sampled pieces came from different bifolios rather than one tiny corner of one leaf. Their compatible results make a simple scenario in which one modern maker assembled the codex from wildly different parchment centuries less attractive. That still does not reveal when every pen stroke was added, but it gives the book a coherent material horizon.
The correct lesson is therefore modest and powerful:
Voynich is not floating outside history. Its parchment belongs to a measurable material world.
The First Discipline: Date the Claim You Are Making
When someone says “the Voynich Manuscript dates to the fifteenth century”, ask what exactly they mean.
- Date of the animal skin?
- Date of parchment manufacture?
- Date of writing?
- Date of drawing?
- Date of colouring?
- Date of first assembly?
- Date of present binding?
- Date of folio numbering?
- Date of a later marginal note?
- Date of a documented owner?
Those can be different dates.
A surprising amount of bad historical reasoning begins when the date of one layer quietly becomes the date of every layer.
For a deeper treatment of this distinction, see The Manuscript Before the Mystery and What We Actually Know.
Part I continues below: inks, pigments, quire engineering, missing leaves, page order, provenance and the difference between an old object and a solved object.
Ink, Paint and the Danger of Asking Materials to Translate
Once the parchment date pushed the physical substrate into the early fifteenth century, the obvious next question was whether the writing and colour behaved like historical manuscript materials too.
McCrone Associates examined and sampled the manuscript in January 2009. Their report describes brownish-black writing of variable darkness, generally consistent in letter size within individual pages, and writing that appears to have been made with a quill pen. Under ultraviolet illumination, the text showed characteristics suggestive of iron-gall ink. Elemental work on selected samples found combinations including carbon, iron, sulphur, potassium and calcium, with additional trace elements in some samples.
The important point is not that a laboratory discovered a magic chemical fingerprint saying “Northern Italy, 1427” or “Bohemia, 1431”. It did not. Materials such as iron-gall ink were used across enormous geographical and chronological ranges.
Materials analysis is strongest when it answers material questions.
- Does the tested ink contain components compatible with historical writing practice?
- Are obvious modern industrial pigments present?
- Do text and drawing inks share microscopic or chemical features?
- Can one visible layer be distinguished from another?
- Do later marks behave differently from the main manuscript?
Those are excellent questions. “What language is this?” is not a chemical question.
This may sound almost comically obvious, but Voynich repeatedly tempts us to promote a legitimate result into a different category of claim. A pigment study becomes a date. A date becomes a location. A location becomes an author. An author becomes a language. A language becomes a translation.
Every arrow in that chain requires new evidence.
What the Colours Tell Us — and What They Do Not
The manuscript is famous in reproduction for green pools, blue washes, reddish details and plant colours that sometimes seem almost cheerful beside the unreadable script. Yet the colouring is one of the easiest layers to over-romanticise.
The drawings are generally outlined in ink and then coloured. The application is often much less refined than the precision of the writing. Some areas are uneven, some washes wander over boundaries, and different intensities have encouraged proposals that colouring occurred in more than one stage or perhaps involved more than one painter.
Those possibilities are worth examining, but a difference in colour density is not automatically a different person. Pigment concentration, water, brush loading, later retouching, deterioration and local working conditions can all change appearance. The manuscript gives us variation; the historical explanation for that variation must still be demonstrated.
More importantly, colour is not a transparent legend. Green does not automatically mean “healthy”, blue does not automatically mean “water”, red does not automatically mean “blood”, and a repeated colour does not prove that two illustrated objects have the same semantic role.
Colour can organise a page without functioning as a modern colour code.
Text, Drawing and Paint May Have Different Production Histories
One of the most useful habits in manuscript study is to stop imagining “the author” doing everything in one sitting.
A historical codex can involve a designer, one or more scribes, one or more illustrators, rubricators, painters, binders, owners, annotators and later repairers. Sometimes one human performs several of those roles. Sometimes a workshop divides them. Sometimes additions are separated by decades or centuries.
Voynich does not currently allow us to assign all those roles with confidence. That uncertainty should remain visible.
What we can study is sequence at local scale. Does writing stop before an illustration? Does text appear to route around a drawing? Does paint cover a line? Does a line cross a fold? Are corrections visible? Do outlines precede colour? Such relationships can reveal workflow even when the words remain unread.
This is a profound idea: production order is another kind of grammar.
If a scribe consistently writes around an already present drawing, that suggests one workflow. If drawings are inserted into reserved spaces, that suggests another. If the page is designed as one integrated field, that suggests a third. We do not need to know whether a plant label means “sage” to learn how text and image were coordinated.
A Folio Is Not a Page in the Modern Sense
To understand the physical Voynich, we need four small pieces of book vocabulary.
Sheet
A piece of parchment before or during folding.
Bifolium
A sheet folded so that it produces two leaves joined at the fold. A bifolium therefore creates four writing surfaces: front and back of one leaf, front and back of the other.
Folio
One leaf. Its front is conventionally called the recto and its back the verso. Thus f1r means folio 1 recto; f1v means folio 1 verso.
Quire or gathering
A set of folded sheets nested together and sewn through their central folds. Several gatherings are then bound to make the codex.
Why should a general reader care?
Because bifolia preserve relationships that page numbers can hide.
Imagine taking one large sheet, folding it, and writing on the resulting leaves. Two pages that end up far apart in normal reading order can belong to the same original physical sheet. Conversely, two pages facing each other in the present binding may have become neighbours only after assembly.
For Voynich, that means a theory based only on “the next page” may be weaker than a theory that respects the actual bifolium structure.
The Manuscript Was Built in Gatherings, Not Downloaded as a PDF
The surviving text block contains 102 folios. A commonly used reconstruction suggests an original sequence of 116 numbered folios, of which fourteen are now absent. The surviving material is organised into eighteen extant gatherings, while the historical quire numbering runs to twenty because quires 16 and 18 are missing.
A fairly standard Voynich gathering can contain four bifolia, producing eight folios. But the manuscript is not mechanically regular. Some gatherings differ in size. Several sheets are extended into foldouts. One exceptionally large sheet opens in both horizontal and vertical directions.
The engineering matters because information was being fitted to material space.
If the creator merely needed another paragraph, a normal page would have sufficed. A foldout suggests that some information benefited from being seen across a larger continuous field. That does not tell us whether the field is a map, a cosmological scheme, a process diagram, a mnemonic structure or something else. It does tell us that the relationship among elements on that surface mattered enough to resist ordinary pagination.
This is especially important when we reach the Rosettes foldout in Part II. Its geometry may be meaningful even if none of its individual motifs can yet be translated.
Foldouts Are Not Decoration
Foldouts change the act of reading.
A reader must stop, open, rotate, expose and sometimes refold the material. The book becomes interactive in the most literal medieval sense: information appears because the reader physically transforms the page.
That gives us another useful distinction:
Reading order and viewing order are not always the same thing.
A long line can cross a fold. A diagram can occupy several panels. A folio number may be written where the corner appears only when the sheet is folded. What looks fragmented in a digital viewer may have been experienced as one unfolding surface by a historical reader.
The digital facsimile is invaluable, but the original object continually reminds us that a scan is a representation of a book, not the book itself.
Fourteen Missing Folios — But Not One Kind of Missingness
“Fourteen pages are missing” sounds simple. It is also imprecise.
The commonly reconstructed missing folios are:
- f12 — apparently cut out; a stub remains.
- f59–f64 — six folios corresponding to three missing bifolia from the centre of quire 8.
- f74 — apparently cut out; cutting evidence is visible around the surviving structure.
- f91–f92 — associated with missing quire 16.
- f97–f98 — associated with missing quire 18.
- f109–f110 — the missing central bifolium of the final gathering, quire 20.
That totals fourteen folios, but the mechanisms are different.
A leaf deliberately cut from a bound book is not the same event as a bifolium disappearing from the centre of a gathering. An entire gathering missing is not the same event as a single leaf being excised. Different loss patterns may imply different moments in the manuscript’s life.
And here is the crucial epistemic rule:
A missing folio can explain why evidence is incomplete. It cannot serve as evidence for whatever we hope the missing folio contained.
If a theory works only by saying “the missing page probably explained this”, then the missing page is functioning as an imagination-shaped plug.
Our separate article The Missing Leaves and Quire Reconstruction goes much further into this problem.
The Numbers Were Added Later
The folio numbers that modern researchers use are extremely helpful. They are not necessarily part of the manuscript’s original information system.
The old foliation runs from 1 to 116, skipping the numbers belonging to leaves now missing. The numbers appear to have been added after the manuscript’s original production, and the foldouts were folded when the numbering was applied. Quire marks form another numbering layer.
These later navigational systems are evidence too. They tell us something about how later custodians encountered, organised and handled the object.
They also produce a warning. When we say “folio 72” we are using a later address. The address is useful; it is not the meaning of the room.
The Present Page Order May Not Be the Intended Original Order
This is one of the most consequential physical facts for interpretation.
Researchers have long noted reasons to suspect that some bifolia or gatherings may not survive in the originally planned sequence. The present sewing is old, but an old binding is not automatically the first binding. Rebinding is ordinary manuscript history.
Suppose a modern scientist found a laboratory notebook whose loose sheets had been rebound two centuries later. Page 30 might now face page 31, yet those sheets might originally have belonged to different experiments. If the handwriting were unreadable, adjacency would become dangerously seductive.
The same caution applies here.
A visual transition across facing pages can be interesting. It is not automatically proof of conceptual sequence. A repeated object on consecutive folios can be meaningful. It is not automatically an operational chain. A theory proposing “first the plant is harvested, then processed, then bathed with, then stored” must prove that those pages belong to a valid sequence rather than inheriting sequence from the current binding.
That single codicological correction eliminates a remarkable amount of false certainty.
Then the Book Enters History
The material evidence brings us into the fifteenth century. The documentary trail does not immediately meet us there.
That gap matters.
There is no surviving fifteenth-century title page saying who commissioned the manuscript. No colophon politely tells us that a named scribe completed it in a named city on a named feast day. No securely identified library catalogue from the year of production has yet been shown to describe this exact codex beyond reasonable doubt.
Instead, the book becomes historically visible in fragments. A faded ownership mark. A set of seventeenth-century letters. A recollection about an emperor. A transfer between scholars. Then silence again. Then a twentieth-century rare-book dealer.
The temptation is to turn those fragments into a smooth story. Good history does the opposite. It keeps the seams visible.
Jacobus Horčický de Tepenec: A Name on the First Folio
One of the most important physical traces on the manuscript is not written in Voynichese at all.
On the first folio is a faint ownership inscription associated with Jacobus Horčický de Tepenec, also known in Latinised form as Jacobus Sinapius. He was a pharmacist and court figure connected with Emperor Rudolf II in Prague. The inscription became difficult to see and has been examined under ultraviolet light; comparisons with ownership inscriptions in other books support its identification.
The form “de Tepenec” is significant because Horčický received that noble title in 1608. If the inscription is his, it therefore belongs after that elevation and before his death in 1622.
That gives us something much stronger than “the pictures look Bohemian”. It is direct physical evidence connecting the manuscript with a person in Rudolfine Prague.
But even here we should resist doing too much.
Ownership in Prague around the early seventeenth century does not prove manufacture in Prague around the early fifteenth century. A book can travel farther than its owner. It can be bought, inherited, exchanged, confiscated, gifted, borrowed or carried across borders.
Custody is evidence of custody. It is not automatically evidence of origin.
Rudolf II and the Famous 600 Ducats
One of the most repeated stories about Voynich is that Holy Roman Emperor Rudolf II bought the manuscript for 600 ducats and believed it to be the work of Roger Bacon.
The story is historically important. The way we phrase it is even more important.
The surviving source is the covering letter that Johannes Marcus Marci sent with the manuscript to Athanasius Kircher in the 1660s. Marci reports information he had heard from Raphael Mnishovsky: that the book had belonged to Rudolf and that the emperor had paid 600 ducats to the person who brought it. Marci also reports the belief that Roger Bacon was the author, while explicitly withholding his own judgment.
That is not the same evidentiary object as a surviving imperial purchase receipt saying: “Paid 600 ducats for this manuscript, now Beinecke MS 408.”
It is reported testimony transmitted through a later letter.
The distinction does not make the Rudolf story worthless. Mnishovsky had court connections, Rudolf was famously interested in learned, esoteric, artistic and scientific material, and the later Horčický inscription places the book in an environment connected with Rudolf’s court. The story fits a plausible historical world.
But plausibility is not a receipt.
This is why our dedicated article calls the problem The Broken Provenance Chain. “Broken” does not mean “nothing is known”. It means the chain contains links of different evidentiary strength.
Roger Bacon: A Theory That History Learned to Let Go
Roger Bacon was irresistible to early twentieth-century Voynich culture.
He was medieval. Brilliant. Associated in popular imagination with secret knowledge. A scholar of language, optics, mathematics and natural philosophy. Exactly the sort of person around whom an unreadable illustrated manuscript could acquire a magnificent legend.
Wilfrid Voynich took the attribution seriously and promoted it. William Romaine Newbold later constructed an elaborate proposed reading that also supported a Bacon connection.
The problem was not that Bacon lacked charisma.
The problem was evidence.
Newbold’s microscopic interpretation did not survive critical examination. More decisively for authorship, radiocarbon dating later placed the parchment in the early fifteenth century, long after Bacon’s death in the thirteenth century.
A historical theory can therefore be culturally powerful, technically ingenious and wrong.
That is not embarrassing. It is the normal price of serious inquiry.
Georg Baresch: The Man Who Admitted the Book Defeated Him
By the 1630s the manuscript was in the orbit of Georg Baresch, a Prague alchemist whose role is far more important than his fame outside Voynich studies.
Baresch did something wonderfully recognisable: he possessed a book he could not read, developed ideas about what it might contain, and sought help from someone reputed to understand difficult scripts.
That someone was Athanasius Kircher, the extraordinarily learned Jesuit polymath in Rome. Kircher worked across languages, antiquity, natural philosophy, mathematics, music, magnetism, Egyptology and other fields. His reputation made him a natural target for anyone holding a mysterious manuscript.
Baresch sent copied material toward Kircher and tried more than once to attract his attention. A surviving 1639 letter is central to the documentary history. Baresch speculated about the manuscript’s intellectual character, including a possible connection with Egyptian knowledge, but speculation by an owner is not the same thing as a medieval title.
This is another subtle point. A seventeenth-century reader’s theory is evidence about seventeenth-century reception. It is not automatically evidence about fifteenth-century authorship.
The distinction is easy to say and surprisingly hard to maintain.
Marci Sends the Sphinx to Kircher
After Baresch, the manuscript passed to Johannes Marcus Marci, a physician, scholar and rector of Charles University in Prague. Marci knew Kircher and eventually sent the manuscript to him in Rome, accompanied by the famous letter that survives with the codex.
That letter is precious because it does several jobs at once.
- It connects a specific historical manuscript transfer to named people.
- It records a previous ownership story about Rudolf II.
- It records the Bacon attribution as something Marci had heard rather than something he established.
- It shows that the text was already unreadable to learned seventeenth-century owners.
- It places the manuscript inside a network of scholars who actively exchanged difficult texts and intellectual problems.
In other words, by the time Voynich becomes historically visible to us in early modern correspondence, it is already a mystery.
That is important. The opacity is not a twentieth-century invention caused by modern scholars forgetting an obvious medieval language. People much closer to the manuscript’s early history were already unable to make ordinary sense of it.
Kircher Did Not Solve It
Kircher’s enormous confidence in deciphering ancient and exotic systems makes this part of the story almost too perfect. If any seventeenth-century polymath was going to look at an unknown script and believe himself capable of penetrating it, Kircher was a strong candidate.
Yet no convincing Kircher decipherment of the Voynich Manuscript survives.
The book appears to have entered Jesuit collections and then largely disappeared from the active scholarly record for centuries.
Silence, however, is not proof of continuous neglect. It means the documentary trail available to us becomes sparse.
The Long Quiet: Rome, the Jesuits and the Problem of Invisible Custody
Historical objects often spend most of their lives doing nothing dramatic.
They sit on shelves.
That is one reason provenance is difficult. An object can remain physically safe for a century while leaving almost no new narrative evidence. The absence of exciting records does not imply the absence of the book.
The manuscript is generally associated with the Jesuit holdings connected to the Collegio Romano and later with material transferred to Villa Mondragone near Frascati. Nineteenth-century political upheaval in Italy affected church property and library arrangements; volumes could be reassigned or sheltered within collections in ways that complicate simple institutional ownership stories.
By the early twentieth century, the manuscript was among books available in a discreet sale from Jesuit holdings.
1912: Wilfrid Voynich Walks In
Wilfrid Michael Voynich was exactly the sort of person a manuscript like this needed in order to become famous.
He was not merely a collector. He was a rare-book dealer with an eye for important material, a talent for building narratives around discoveries and access to scholarly networks that could examine what he found.
In 1912 he acquired a group of manuscripts from the Jesuit collection at Villa Mondragone near Rome. Among them was the unreadable codex.
From that point, the object entered its modern life.
The name changed in practical effect. Instead of being an anonymous manuscript in a religious collection, it became the Voynich Manuscript: an object attached to the dealer who brought it to public scholarly attention.
That naming is historically convenient and intellectually misleading in one useful way. Voynich did not make the book merely because the book now carries his name.
The early-fifteenth-century parchment and seventeenth-century documentary trail make a simple 1912 manufacture scenario untenable.
The Modern Hoax Question
Could Wilfrid Voynich have created the manuscript himself to sell an exciting mystery?
As a modern-forgery hypothesis, that has severe problems.
- The parchment dates centuries before Voynich.
- The manuscript is connected with seventeenth-century correspondence.
- The Horčický ownership mark ties the object to an early modern Central European context before Voynich’s lifetime.
- The tested materials are consistent with historical manuscript materials rather than presenting a simple suite of modern anachronisms.
But notice what eliminating a modern Voynich forgery does not eliminate.
It does not prove that the visible text encodes ordinary prose. A fifteenth-century person could create a cipher, an artificial script, an abbreviation system, a mnemonic notation, pseudo-text, a constructed language, a specialised technical notation or something for which our categories are incomplete.
“Old” and “meaningful in the way we expect” are separate claims.
Ethel Voynich, Anne Nill and Hans P. Kraus
After Wilfrid Voynich died in 1930, the manuscript remained within his personal and professional circle. His wife, Ethel Lilian Voynich—herself a notable novelist and the daughter of mathematician George Boole—eventually left the manuscript to Anne Nill, a close friend and former assistant.
Nill sold it in 1961 to the major rare-book dealer Hans P. Kraus. Kraus hoped to sell the manuscript at a high price but did not find a buyer willing to meet his expectations.
In 1969 he donated it to Yale University.
There the anonymous book acquired its modern institutional identity: Beinecke MS 408.
Why Yale Changes the Story
Institutional custody does not solve a manuscript. It changes what can be done with it.
At the Beinecke, the Voynich Manuscript became part of a research library rather than a dealer’s stock. Conservation, cataloguing, photography, scientific sampling, controlled access and eventually high-quality digital reproduction made forms of study possible that earlier owners could not have imagined.
The modern Voynich research community is therefore built on a paradox.
The manuscript remains unread, but it has never been more observable.
A student in Singapore can inspect a digital folio that seventeenth-century scholars might have crossed Europe to see. Researchers can compare transcriptions across thousands of tokens. Material scientists can analyse microscopic samples. Palaeographers can compare handwriting. Codicologists can reconstruct gatherings. Linguists can measure entropy and positional effects. Computer scientists can test models against the entire corpus rather than a handful of reproductions.
More access has not made the answer obvious.
It has made weak answers easier to challenge.
Provenance Is a Chain, Not a Vibe
Voynich research frequently moves from resemblance to geography too quickly.
A plant illustration resembles an Italian herbal, therefore Italy. A zodiac label resembles a Romance month name, therefore a particular region. A diagram resembles a German cosmological manuscript, therefore Germany. A bathing scene resembles a Central European balneological tradition, therefore Bohemia.
Each comparison may be worth investigating.
None of them, by resemblance alone, creates provenance.
A provenance claim needs a mechanism connecting the actual object through time: ownership marks, inventories, letters, bindings, library records, purchase documentation, known transfers, material histories or some other evidence capable of tracking this codex rather than merely identifying a similar cultural feature.
This is especially important for eduKateSG’s comparisons with the Carrara Herbal, Padua medical manuscripts, Masson 116, Sloane 4016, Egerton 747 and other medieval scientific or medical books. Those comparisons can show what kinds of information architecture were historically possible. They can illuminate how image, label, plant, preparation, zodiac, medicine or diagram interacted in other manuscripts.
They do not automatically make any of those manuscripts an ancestor of Voynich.
Comparator is not source. Resemblance is not descent. Plausibility is not provenance.
The Cover Is Not the Beginning
Modern books train us to trust the cover. It usually belongs to the same publication event as the text. It tells us the title, author, publisher and often the intended subject before we reach page one.
Voynich gives us no such luxury.
The present cover and binding belong to the later history of the object rather than functioning as an original title-bearing wrapper supplied by the creator. There is no surviving original front board saying what the book was called. No medieval spine title. No author name. No table of contents that settles the six familiar modern “sections”.
That absence changes how we should speak.
“Herbal section”, “astronomical section”, “biological section”, “cosmological section”, “pharmaceutical section” and “recipe section” are useful modern descriptive labels. They help us navigate the visual changes in the manuscript.
They are not recovered chapter titles.
The difference is essential. A page with a plant-like drawing can reasonably be called herbal-looking. That does not mean we have proved that its text is a botanical description. A page with a circular arrangement and zodiac-like figures can reasonably be grouped with astronomical or astrological material. That does not mean the text is a horoscope. A page containing containers and detached plant parts can resemble pharmaceutical imagery without being a prescription manual.
Classification is useful before decipherment. It becomes dangerous when the classification quietly pretends to be decipherment.
Six Familiar Sections, Zero Proven Chapter Titles
For navigation, researchers commonly describe the manuscript through several broad visual regimes.
Plant-dominated pages
Many early folios present a large plant-like figure with surrounding text. Some plants appear recognisable in parts; others seem composite, distorted or simply difficult to identify. The temptation to assign modern species names is powerful, but botanical resemblance is often non-unique, especially after generations of copying, stylisation and imperfect preservation.
Circular and zodiac-like pages
Other folios contain circular arrangements, stars, radial text and figures associated with zodiac signs. Some later month-name annotations are written in a different, more familiar script. The circular regime is clearly different from ordinary plant pages, but “astrological” remains a functional interpretation rather than a translation of the main text.
Human figures, pools and tubes
A striking group of pages, especially in Quire 13, contains numerous small human figures—often interpreted as female—within pools, vessels or interconnected tube-like structures. The pages have been compared with bathing, anatomy, generation, cosmology, medicine, water systems and symbolic diagrams. The images are real. The category “balneological” is a useful comparison, not a secure caption supplied by the manuscript.
Large cosmological-looking diagrams
Foldouts and circular structures create some of the manuscript’s most spectacular pages, including the Rosettes foldout. Their relational geometry is undeniable. Their precise referents are not.
Containers and detached plant parts
Later pages show smaller plant components and rows of elaborate vessel-like forms. Because comparable historical manuscripts can connect ingredients, containers and medicine, the “pharmaceutical” label is understandable. Yet a vessel shape does not decode the adjacent label.
Starred short paragraphs
The final major text regime contains many short entries, frequently marked by star-like symbols. Because this resembles the segmentation of recipes or records in other manuscripts, it is often called the recipe section. The strongest immediate evidence, however, is simply that the text has been deliberately segmented into many addressable units.
Part II will walk through these worlds in detail. For now, the important object-level conclusion is this:
The manuscript changes visual and textual regimes, but our modern names for those regimes are hypotheses about function, not recovered medieval labels.
What the Object Rules Out
A great mystery can make knowledge feel smaller than it is. We say “unsolved” and accidentally imagine that every theory remains equally possible.
They do not.
The physical object removes real possibilities.
- Roger Bacon as the original writer on these leaves is incompatible with the radiocarbon age of the parchment.
- A straightforward twentieth-century fabrication by Wilfrid Voynich is incompatible with the manuscript’s pre-Voynich historical trail and early material date.
- A theory requiring all surviving pages to remain in guaranteed original conceptual order ignores codicological uncertainty and evidence of loss or reassembly.
- A theory requiring the fourteen missing folios all to have vanished in one event ignores the different physical forms of loss.
- A theory treating later folio numbers, month names or ownership marks as part of the original Voynich script collapses historical layers that can be distinguished palaeographically and contextually.
- A theory that needs one visual resemblance to establish birthplace or direct copying demands more from iconography than iconography can provide alone.
Negative knowledge is still knowledge.
If a thousand doors once seemed open and evidence closes eight hundred, we have not failed because the remaining two hundred are still difficult. We have learned the architecture of the problem.
What the Object Does Not Settle
At the same time, physical evidence has limits.
- It does not identify the main script as an alphabet, syllabary, cipher alphabet, abbreviation system or something else.
- It does not tell us whether spaces correspond to linguistic words.
- It does not identify the underlying language, if a natural language exists underneath the visible system.
- It does not prove that the text is encrypted.
- It does not prove that the text is unencrypted.
- It does not tell us whether the plants are intended as literal botanical portraits, mnemonic composites or transformed exemplars.
- It does not establish that pools represent baths, anatomy, alchemy, cosmology or hydraulic processes.
- It does not identify the Rosettes foldout as a map, city plan, cosmogram, memory system or process diagram.
- It does not tell us the names of the scribes or illustrators.
- It does not give us the exact place of manufacture.
This is not disappointing. It is simply the boundary between material history and semantic interpretation.
An Old Manuscript Can Still Contain Artificial Text
One of the most persistent logical mistakes in Voynich discussion is:
The manuscript is genuinely old; therefore the text must encode an ordinary meaningful language.
The conclusion does not follow.
Historical Europe contained ciphers, secret alphabets, shorthand, mnemonic devices, magical scripts, alchemical notation, astronomical symbols, scribal abbreviation systems, invented signs, tables, diagrams and texts created for audiences very different from us.
A medieval or Renaissance artefact can be authentic as an artefact while its visible text is highly transformed, convention-bound, deliberately opaque or even partly non-semantic.
The age of the object therefore authenticates its historical existence. It does not pre-decipher its content.
A Meaningful Text Can Still Look Artificial
The opposite mistake is just as common.
Voynichese has unusual repetition, strong positional preferences, restricted combinations and low-level predictability that can make it feel “generated” or “too regular” to some observers.
But transformed meaningful systems can also acquire unusual surface statistics. Abbreviation, substitution, null insertion, syllabic encoding, constrained orthography or other processes can change what a language looks like.
This is why Part III will not ask whether Voynichese “looks like language” as if visual intuition could settle the matter. It will ask what measurable constraints exist, what mechanisms can reproduce them, and what those mechanisms predict elsewhere in the manuscript.
The Book Is Damaged, But It Is Not a Ruin
Because missing leaves receive so much attention, it is easy to imagine Voynich as a shattered fragment.
That overstates the damage.
More than two hundred writing surfaces survive, including major continuous runs, elaborate foldouts and the large final gathering. Enough text remains for detailed statistical analysis. Enough physical structure remains for codicology. Enough visual material survives to compare recurring motifs across the book.
The missing material matters enormously for reconstruction, but the surviving codex is not a handful of disconnected scraps.
That creates an uncomfortable implication for grand decipherment theories: there is enough surviving material that a strong theory should explain more than one lucky paragraph.
A Single Folio Is Not the Manuscript
Voynich “solutions” often begin with one spectacular success.
A plant is identified. A word is read. A zodiac sign is matched. A diagram is recognised. A short phrase can be massaged into Latin, Hebrew, Romance, Germanic, Turkic, Semitic or another language.
That can be a legitimate beginning.
It cannot be the end.
A manuscript-scale explanation must survive manuscript-scale diversity. It has to cope with plant pages and labels, circular pages, foldouts, Currier differences, line-position effects, repeated word families, star-marked entries, local vocabulary and scribal variation. It has to explain not only the place where it succeeds but the places where its mechanism should have worked and did not.
We return to that standard in Part V and in What a Real Voynich Decipherment Must Survive.
How to Look at a Voynich Folio Without Fooling Yourself
Try this sequence.
1. Describe before interpreting
Say what is physically present. “A circular figure with twelve repeated outer units” is safer than “a zodiac calendar” until the relevant features have been established.
2. Separate main script from later additions
Not every visible mark belongs to the same writing event. Later numbers, month names, ownership inscriptions and marginal writing can be historically valuable precisely because they are later.
3. Ask what physical sheet you are looking at
Which bifolium? Which gathering? Is it a foldout? What lies on the other side? Does the theory depend on current adjacency?
4. Notice relationships, not only objects
Where does text begin and stop relative to the image? Which labels are close to which elements? Do repeated forms occupy repeated positions? Does the page have a centre, sequence, hierarchy or symmetry?
5. Test the comparison elsewhere
If one vessel shape means one process, what happens when that vessel recurs? If one glyph is a letter, does the reading survive other positions? If one plant part has a semantic value, does that value predict anything on another page?
6. Record the failure
The point where an explanation stops working is not useless. It tells us the boundary of the explanation.
This last step is why the failed ideas in our broader Voynich work matter. A theory that fails cleanly can leave behind a better map of the manuscript than a theory that survives only by changing its rules whenever it meets resistance.
Why “Everything Is Connected” Is Not Enough
Voynich genuinely contains recurring forms. Plants reappear in different scales. Stars recur. Circular structures recur. Similar token families cluster. Labels share patterns with running text while also behaving differently. Visual motifs sometimes seem to migrate from one regime to another.
It is therefore reasonable to search for a connected system.
But connection itself is cheap.
In a manuscript created by one workshop, many features can be connected simply because the same people, pens, conventions and materials recur. A repeated shape can be a semantic symbol, a scribal habit, a decorative convention, a structural marker, a copied exemplar feature or an artefact of how the page was built.
The hard question is not:
Can these two things be connected?
With enough imagination, almost anything can.
The hard question is:
Does the proposed connection make a new prediction that survives somewhere else?
That is where a pattern becomes evidence.
What We Know at the End of Part I
Before touching the language, we can already say quite a lot.
- Voynich is a real historical codex, not a purely modern internet puzzle.
- Its parchment belongs to the early fifteenth century according to radiocarbon measurements from multiple samples.
- The parchment is calfskin.
- The writing and drawing materials examined scientifically are compatible with historical manuscript production.
- The book was physically engineered from bifolia, gatherings and unusual foldouts.
- The surviving book is incomplete.
- The missing material disappeared in more than one physical pattern.
- The current foliation is later than the main production.
- The present sequence cannot automatically be assumed to preserve every original conceptual relationship.
- The manuscript was in Prague-linked ownership by the early seventeenth century.
- The Rudolf II purchase story is historically important but reaches us through reported testimony.
- The Roger Bacon authorship claim fails against the material date.
- Baresch, Marci and Kircher show that the manuscript was already difficult to interpret in the seventeenth century.
- Wilfrid Voynich acquired it in 1912; later it passed through Ethel Voynich, Anne Nill and Hans P. Kraus before reaching Yale in 1969.
- Authenticity of the object does not prove ordinary linguistic plaintext.
- Unusual text statistics do not prove meaningless generation.
- Comparative manuscripts can establish historical possibility without establishing direct ancestry.
That is a surprisingly solid floor.
Now we can finally open the book.
Part I Field Guide: Questions a Careful Reader Should Be Able to Answer
A master article should not merely leave the reader impressed. It should leave the reader better equipped.
So before we leave the physical manuscript and enter its illustrated world, here are the questions that most often become tangled together—and the cleanest answers the evidence presently allows.
Why is it called the Voynich Manuscript?
Because Wilfrid Michael Voynich acquired the codex in 1912 and brought it into modern scholarly and public attention. The name tells us about the manuscript’s modern rediscovery, not its fifteenth-century title. We do not know what its makers called it, or whether it ever possessed a title in the modern sense.
What is its official library name?
At Yale it is catalogued as Beinecke MS 408, in the Beinecke Rare Book and Manuscript Library. “Voynich Manuscript” is the familiar name; MS 408 is its modern institutional call number.
How old is it?
The safest concise answer is: its sampled parchment dates to the early fifteenth century. Four radiocarbon samples produced mutually compatible results, commonly summarised as 1404–1438 at the reported 95% confidence level. That establishes a material horizon. It does not by itself give an exact year for the writing, drawings, colouring or assembly.
Does carbon dating prove the ink is from the same period?
No. Radiocarbon testing was performed on parchment. Scientific examination of inks and pigments found materials compatible with historical manuscript manufacture, but the ink was not directly radiocarbon-dated into the same narrow window. The parchment date and the material compatibility support one another without becoming identical measurements.
Is it written on paper?
No. It is written on parchment—processed animal skin. Protein analysis identified calfskin. That matters because the manufacture, folding, sewing and survival of parchment preserve clues about the object that a purely textual analysis would miss.
How many pages survive?
The surviving codex contains about 102 folios, meaning roughly 204 ordinary recto-and-verso writing surfaces before accounting for the unusual geometry of foldouts. Older foliation indicates a sequence reaching 116, with fourteen folio numbers now absent. “About 240 pages”, “204 pages” and similar figures sometimes appear in general accounts because people count foldouts, surfaces or numbered leaves differently. For serious work, folio notation is safer than a casual modern page count.
Are the missing folios all mysterious removals?
No. Their physical situations differ. Some leaves appear to have been cut out; others belong to lost bifolia; two numbered gatherings are missing. The manuscript does not give us one simple “someone removed fourteen secret pages” event. Its losses have structure.
Can the missing pages contain the key?
They can contain anything historically possible, which is exactly why they cannot be used as positive evidence. Perhaps a missing folio once contained a title, a diagram, an index or nothing unusually helpful. We do not know. A responsible theory must work with surviving evidence rather than borrowing certainty from absent material.
Is the current page order original?
Not necessarily. The physical structure preserves important bifolium and quire relationships, but rebinding, lost material and later foliation mean that present adjacency cannot automatically be equated with original conceptual sequence. This does not make all order meaningless. It means order has to be reconstructed rather than assumed.
Why do some pages unfold?
Because the makers used larger or extended surfaces when ordinary folio geometry was insufficient. Foldouts allow diagrams and relationships to be viewed across a continuous field. Their existence is therefore evidence about information architecture even though we cannot yet name the exact information represented.
Was it made in Italy?
Possibly. Was it made in Central Europe? Also possible. Various palaeographic, iconographic and historical arguments have been advanced for regions within medieval Europe, and comparisons with Italian, Germanic, Bohemian and other manuscript traditions can be illuminating. But the exact place of manufacture remains unresolved. “European” is safer than forcing a city from one attractive resemblance.
Did Rudolf II really own it?
The Rudolf connection is historically plausible and important, but the famous purchase story reaches us through later testimony recorded by Marci. A physical ownership mark connects the manuscript more directly with Jacobus Horčický de Tepenec, a figure associated with Rudolf’s court. The prudent formulation is therefore that the Rudolfine court forms an important part of the manuscript’s early modern provenance story while the exact transaction remains less directly documented than a surviving receipt would be.
Did Rudolf pay 600 ducats?
That amount is reported in Marci’s letter from information attributed to Raphael Mnishovsky. It is evidence that such a story circulated among people connected with the manuscript. It is not the same category of evidence as an imperial account-book entry securely identifying MS 408.
Was Roger Bacon the author?
Not of the surviving codex. Bacon died in the thirteenth century, while the parchment dates to the early fifteenth. A later manuscript could theoretically copy a lost Baconian work, but that is a completely different and much weaker claim requiring independent evidence. The old direct-authorship theory does not survive the dating.
Was John Dee involved?
John Dee has appeared in Voynich provenance speculation because he moved in relevant learned and courtly worlds and visited Rudolf’s Prague. Yet no secure chain identifies Dee as the manuscript’s seller or owner. A historically convenient person is not automatically the missing person in a provenance gap.
Could Wilfrid Voynich have forged it?
A straightforward modern fabrication by Voynich is contradicted by the early material date and pre-1912 documentary and ownership evidence. More elaborate scenarios involving old blank parchment would still have to explain the seventeenth-century documentary trail and other physical layers. The modern-forgery theory no longer provides the economical explanation it once might have seemed to offer.
Does that prove every mark in the book is fifteenth-century?
No. Later foliation, ownership marks, month names and marginal writing show that the manuscript accumulated additional marks during its life. Historical authenticity is layered.
Why do people call parts of it herbal, astrological, biological or pharmaceutical?
Because those are useful visual comparisons. Large plants resemble herbal manuscript conventions. Zodiac-like figures and radial diagrams resemble astronomical or astrological material. Pools and human figures invite balneological comparison. Vessels and plant parts invite pharmaceutical comparison. Star-marked short entries resemble recipe or record segmentation. The names help us navigate; they do not translate the text.
Do the pictures prove the subject?
No. Pictures constrain interpretation, sometimes strongly, but the same visual object can serve multiple functions. A star can be astronomical, decorative, enumerative or structural. A vessel can represent storage, preparation, classification or something symbolic. A human figure in liquid does not come with a modern caption saying “therapeutic bath”.
Why has nobody simply recognised the language?
Because the problem may not be “identify this alphabet and read the words”. The visible glyphs could encode language through an unfamiliar orthography, abbreviation scheme, cipher, syllabic system, transformation or mixed mechanism. Even the spaces may not delimit ordinary words. The fact that no conventional reading has won acceptance is itself evidence that the representation layer matters.
Is it definitely encrypted?
No. “Cipher manuscript” is a historical and cataloguing label, not a demonstrated mechanism. Cryptographic hypotheses are important because the text is opaque and structurally constrained, but encryption remains one family of explanations among several.
Is it definitely meaningful?
That is also unproved. The text is highly structured, non-random in ordinary senses, locally patterned and rich in distributional regularities. Those properties require explanation. They are compatible with meaningful language under transformation, but carefully generated non-semantic or partly semantic systems can also produce surprisingly rich structure. The correct target is mechanism, not intuition.
Has AI solved it?
No publicly demonstrated AI solution has satisfied the manuscript-scale evidentiary burden. Machine learning can classify pages, compare distributions, identify clusters, test candidate languages, measure glyph behaviour and discover patterns humans might overlook. None of that permits a model to rename pattern recognition as translation. A real solution must be reproducible, predictive and capable of surviving unseen text and the manuscript’s multiple regimes.
Why is Voynich still worth studying if we cannot read it?
Because unreadability is only one property of the manuscript. It is also a fifteenth-century material object, an exercise in codex engineering, a large structured dataset, a history of scholarly failure, a case study in inference, an archive of visual relationships and an unusually demanding test of how humans distinguish evidence from resemblance.
In that sense, Voynich is not merely an unsolved message. It is a laboratory for thinking.
Part I Reading Map: Go Deeper Without Losing the Master
The master article is designed to remain readable from beginning to end. The satellite articles are where individual problems can be expanded without turning every paragraph here into a monograph.
- What We Actually Know — the compact evidence-first doorway.
- The Manuscript Before the Mystery — parchment, codex and production before interpretation.
- The Broken Provenance Chain — Rudolf II, Horčický, Baresch, Marci, Kircher, Voynich and the difference between reported and direct evidence.
- The Missing Leaves and Quire Reconstruction — what physical loss can and cannot tell us.
- Marginalia and the Later Hands — separating the main production from later use and annotation.
- The Scribe Problem — whether visible hand variation implies multiple scribes and what that means for production.
- When Information Survives but Knowledge Dies — the larger problem of intact representation after context has disappeared.
- The Carrara Herbal — a historically grounded comparator showing what image, language and medical knowledge can do together without being declared a Voynich source.
- Masson 116 — a useful comparator for image-led information and incomplete textual context.
Primary Sources and Research Foundations for Part I
Voynich has accumulated an enormous secondary literature. For the physical object, we privilege sources closest to the manuscript and to the measurements being discussed.
Yale Beinecke Rare Book and Manuscript Library
The Beinecke Cipher (Voynich) Manuscript is Yale’s public collection guide. It provides the institutional description, access information, broad provenance account, links to manuscript images and links to scientific examination.
Yale’s public description is an essential starting point, but even an institutional summary should be read alongside the underlying documents. Where a provenance statement derives from a historical report rather than a surviving transaction record, this master preserves that distinction.
McCrone Associates materials analysis
McCrone Associates: Voynich Manuscript analysis report documents the 2009 microscopic and elemental examination of inks and pigments. This is the place to check what the laboratory actually observed rather than relying on simplified claims such as “science proved the ink date”.
Radiocarbon dating and codicological synthesis
René Zandbergen’s radiocarbon documentation assembles the dating results, calibration issues and sample information in a form useful to researchers. His broader origin and material-history pages and quire-layout reconstruction are valuable for understanding bifolia, gatherings, foldouts and missing material.
Zandbergen’s work is a research synthesis rather than the manuscript itself. Its value lies in careful collation of scans, transcriptions, archival material, physical observations and prior scholarship. As everywhere in this master, useful reconstruction remains distinguishable from direct physical observation.
Yale University Press facsimile
The Voynich Manuscript, edited by Raymond Clemens, reproduces the manuscript with scholarly essays on its history, materials and interpretive problems. A facsimile is especially useful because foldouts, sequence and page relationships are difficult to appreciate when the book is encountered only as isolated internet images.
A Note About Sources, Claims and Certainty
There is no single source called “the answer to Voynich”.
A library catalogue knows certain things very well. A materials laboratory knows other things very well. A palaeographer can identify handwriting features that a chemist is not trying to answer. A codicologist can reconstruct sheets that a cryptanalyst may accidentally treat as ordinary page order. A linguist can measure token distribution without knowing what a pigment contains. A historian can assess the provenance of a seventeenth-century report without decoding the script.
World-class research does not flatten those disciplines into one authority. It asks each source to answer the question it is actually equipped to answer.
This is why we will sometimes disagree with the confidence level of a convenient public summary while still using it as a valuable source. The question is not whether a source is “good” or “bad”. The question is whether a particular sentence is supported by the kind of evidence that sentence requires.
What Part I Has Changed
At the beginning of this article, Voynich could still feel like one giant question mark.
It should not feel that way now.
The question mark has acquired edges.
We know the substrate belongs to an early-fifteenth-century material horizon. We know the book is engineered from parchment bifolia and gatherings. We know some leaves are gone and we can distinguish several forms of loss. We know later hands added navigational and ownership layers. We know the book moved through an early modern Central European intellectual world before entering Jesuit custody and, much later, the modern rare-book trade. We know that a direct Roger Bacon authorship fails. We know that a simple Voynich-era fabrication fails. We know that the object’s authenticity cannot decide whether its script encodes ordinary language. We know that a foldout changes how information can be organised. We know that current adjacency can mislead. We know that comparisons can establish possibility without creating provenance.
Those are not decorative facts around the mystery.
They are the walls within which every future solution must fit.
A good Voynich theory does not begin by asking how much of the manuscript it can explain. It begins by asking how much of the manuscript it is forbidden to contradict.
From Object to World
Now we are ready for the most immediate seduction in the manuscript.
The pictures.
They seem easier than the writing. A leaf looks like a leaf. A star looks like a star. A ram looks like Aries. A woman stands in green liquid. A vessel looks as though something could be stored inside it. Nine large circles connect across an extraordinary foldout.
Surely, we think, if the words will not talk, the pictures will.
Sometimes they do.
But pictures have grammar too, and resemblance is not the same thing as meaning.
Part II — The World Inside the Book begins there: with plants that almost identify themselves, circles that almost become cosmology, bodies that almost become medicine, vessels that almost become pharmacy, and a foldout large enough to tempt almost anyone into drawing a map across it.
The word to keep with us is almost.
Part II — The World Inside the Book
If Part I asked us to stop treating Voynich as a disembodied code, Part II asks us to stop treating its pictures as subtitles.
The pictures are not subtitles.
They are evidence, but evidence of a different kind.
A reader opening the manuscript for the first time usually has one immediate advantage over the cryptanalyst: the pictures appear to be recognisable. We see roots, leaves, flowers, circles, stars, human bodies, containers, tubes, pools and something like zodiac signs. The writing refuses us. The images seem to welcome us in.
That welcome is genuine. It is also dangerous.
Visual resemblance produces confidence much faster than it produces proof. We recognise a familiar outline and our mind supplies the missing noun. Then the noun supplies a function. The function supplies a story. By the time we notice what happened, one curved line has become a medical instrument, one pool has become a therapeutic bath, one castle has become a specific Italian city and one plant root has become a known drug.
Voynich teaches an uncomfortable visual lesson:
Seeing what something resembles is not the same as knowing what role it plays.
Quick Read: The Visual World
- The manuscript’s illustrations genuinely change by region, which is why researchers use conventional labels such as herbal, astronomical/astrological, biological/balneological, cosmological, pharmaceutical and recipes/stars.
- Those labels are navigational conveniences, not recovered medieval chapter headings.
- The plant pages resemble European herbal traditions in page architecture, but no complete catalogue of plant identifications has won broad acceptance.
- The zodiac pages contain recognisable signs and later Romance-language month names, but the central Voynichese remains unread.
- The Rosettes foldout is a large relational diagram of nine major circular structures connected across an unfolded surface; “map”, “cosmogram” and “process diagram” remain interpretations.
- Quire 13 contains human figures, pools, tubes and vessels in integrated arrangements. Historical bathing and medical traditions make balneological comparison plausible without proving that the pages are a bath manual.
- The so-called pharmaceutical pages reduce many plants to smaller fragments and place them beside ornate containers and labels, suggesting a change in representational scale or information function.
- The final starred pages largely abandon figurative illustration but retain strong segmentation: hundreds of star-like marginal markers divide text into addressable units.
- Some visual forms recur across different regimes. Those recurrences are more useful when treated as testable relationships than as instant translations.
The Six Sections Are a Map We Drew
Yale’s public description follows the familiar convention of dividing the manuscript according to what its illustrations appear to show: plants; astronomical and astrological material; baths and bathing; cosmological medallions; herbs, roots and containers; and mostly unillustrated text marked by star-like forms.
That convention is extremely useful.
But it is ours.
There is no surviving Voynich table of contents saying:
- Chapter One: Botany
- Chapter Two: Astrology
- Chapter Three: Balneology
- Chapter Four: Cosmology
- Chapter Five: Pharmacy
- Chapter Six: Recipes
The illustrations justify grouping similar pages. They do not tell us what the original maker called the groups, whether the groups had the same intellectual boundaries, or whether the manuscript’s underlying text changes subject exactly when the pictures change.
This is more than a terminological nicety.
If we call a set of pages “pharmaceutical” long enough, a jar begins to look like medicine even before we test whether its labels behave like ingredient names. If we call another set “recipes”, every short paragraph begins to feel procedural. The name starts solving the page for us.
So throughout Part II we will use the conventional labels, but keep a small mental asterisk beside each one.
The Herbal Pages: The Easiest Place to Become Overconfident
Open the manuscript near the beginning and the visual grammar seems almost reassuring.
A plant dominates the page. Text occupies the available space around it. The basic arrangement looks familiar from European herbals: specimen plus writing.
For a moment, Voynich looks normal.
Then you try to name the plant.
Some leaves look plausible. Some roots look surprisingly specific. Some flowers resemble familiar genera. Then the same plant seems to contain parts that point in different botanical directions. A root may look convincing while the flower does not. Leaves may be arranged in an unnatural geometry. A stem may loop in ways that living plants rarely do. Parts can appear exaggerated, rotated, simplified or combined.
This does not make the drawings meaningless. It means “photographic species portrait” is only one model of what a medieval plant image can be.
A Medieval Herbal Is Not a Field Guide
Modern natural-history illustration trains us to expect diagnostic precision. We want leaf margins, venation, flower structure, fruit, scale and habit to line up with a species in a taxonomic system.
Medieval manuscript images could serve different jobs.
- They could help recognise a useful plant.
- They could preserve an inherited pictorial tradition copied from earlier exemplars.
- They could emphasise the part considered medically important.
- They could function mnemonically rather than taxonomically.
- They could compress several growth stages or plant features into one image.
- They could be copied by someone who had never seen the living specimen.
- They could accumulate distortion across generations of copying.
Once those possibilities are admitted, a strange Voynich plant stops forcing us into the crude choice between “accurate species” and “invented fantasy”.
There is a very large middle territory called manuscript transmission.
Why Plant Identification Is Harder Than It Looks
Suppose a Voynich plant has a root like Plant A, a leaf like Plant B and a flower like Plant C.
Several explanations remain possible.
- The illustrator made errors.
- The plant is stylised.
- The plant is composite.
- We are misreading the drawn morphology.
- The species has changed in appearance under domestication or environmental conditions.
- The image descends from a copy tradition whose distortions accumulated.
- The drawing is not intended to represent a biological species at all.
Now add another problem: humans are excellent at resemblance matching.
Given a sufficiently large botanical catalogue, a sufficiently motivated researcher can usually find something that resembles one feature. The correct question is not “Can I find a plant that looks similar?” It is “Does the same identification explain enough independent features better than competing identifications, and does anything else in the manuscript support it?”
A plant identification becomes more useful if it predicts the adjacent label, the surrounding text, a detached fragment elsewhere, a medical use, a geographic context or some repeated internal relationship.
Otherwise it remains resemblance.
The Page Layout Is More Reliable Than the Species Name
Even when we cannot identify a plant, we can observe what the page does.
In many herbal folios, the large drawing occupies the centre or a major vertical zone while paragraphs are fitted around it. Text may stop at the plant, continue on the other side, or occupy open spaces above and below. Some pages contain short labels in addition to running text. This means image and writing were composed as a shared surface rather than merely printed independently and pasted together conceptually.
The page therefore has at least two simultaneous structures:
- object structure — root, stem, leaf, flower, fruit-like or bulb-like forms;
- document structure — paragraphs, labels, interruptions, margins, line endings and available writing space.
Those structures interact whether or not the plant is belladonna, Hypericum, a copied composite or something else entirely.
One of the Most Important Internal Clues: Plants Change Scale
The large herbal drawings are not the only plant-like forms in the book.
Later, in pages traditionally called pharmaceutical, smaller plant fragments appear in rows beside labels and elaborate containers. Researchers have noted internal visual similarities between some of these small fragments and earlier large herbal drawings. Individual comparisons remain debatable, but the broader structural fact is harder to dismiss: the manuscript repeatedly represents plant-like material at different scales and in different page architectures.
This is potentially more informative than forcing a species name.
If a whole-plant page and a later fragment page really refer to the same internal object class, then the manuscript may be changing what it wants the reader to do with that object.
Whole organism → selected part.
Recognition → handling.
Description → classification.
Plant → ingredient.
Those arrows are hypotheses, not translations. But they are testable hypotheses because they predict relationships across sections.
Our dedicated Herbal Pages article explores this identification problem in depth, while What the Pictures Can and Cannot Tell Us establishes the larger visual-evidence rules.
Then the Book Turns in Circles
At some point, the manuscript stops asking us to look at one dominant plant and starts asking us to look around.
The page becomes radial.
Circles nest inside circles. Stars repeat. Small human figures occupy rings. Labels sit beside them. Zodiac signs appear at centres. Text may run around a perimeter instead of across a normal left-to-right line.
The change is not cosmetic. Circular layout changes what counts as “next”. In ordinary prose, the reader expects sequence: line one, line two, paragraph one, paragraph two. In a ring, adjacency can be radial, angular, concentric or cyclic. A centre can govern an outer band. Repetition can encode count. Orientation can matter.
In other words, the page stops behaving like a sheet of prose and begins behaving like an interface.
The Zodiac Signs Are Among the Least Ambiguous Pictures in Voynich
Some Voynich images are difficult to classify at all. The zodiac sequence is different.
Fish for Pisces, a ram for Aries, a bull for Taurus and other familiar zodiac emblems provide unusually strong iconographic anchors. We are not inferring “zodiac” merely because the pages contain circles and stars. The central figures themselves belong to a recognisable European zodiacal tradition.
That makes the zodiac pages extremely valuable.
They are one of the few places where an external semantic class—zodiac sign—can be assigned with much higher confidence than we can assign a species to most herbal drawings.
But even here, one solved picture does not solve its labels.
A Zodiac Sign Gives Us a Context, Not a Translation
If the central emblem is Taurus, we know something about the visual context of that page. We still do not know what the dozens of surrounding Voynichese labels mean.
They could name stars. Degrees. Days. People. Qualities. Subdivisions. Ritual categories. Astrological conditions. Something mnemonic. Something we have not considered.
The ring architecture constrains the options because the labels are attached to repeated positions and repeated figures. Yet an external zodiacal frame remains a frame, not the contents of every box inside it.
This is exactly the kind of place where a future decipherment should become predictive. If a proposed reading says that each label names a day or a degree, the full ring should behave accordingly. Counts, repetitions and relationships should align systematically rather than only where a convenient example can be found.
The Thirty-Figure Problem
Many zodiac diagrams contain around thirty small human figures arranged in rings, often holding star-like objects. The count naturally invites calendrical interpretations because thirty approximates the number of days in a month.
That is a reasonable hypothesis.
It is not self-proving.
The structure may indeed encode days, degrees or another thirty-part division. But a repeated count can arise from multiple systems. The responsible route is to ask whether the labels, ordering, sign transitions and exceptions behave in a way uniquely expected by the calendrical hypothesis.
Voynich rewards counting. It punishes stopping at the first attractive number.
The Month Names Are Real — and Later
The zodiac pages contain something precious: familiar writing.
Month names such as forms corresponding to March, April, May and later months were written near the zodiac figures in a script different from the main Voynichese. Scholars have debated the precise Romance language or dialect reflected in those forms, with Northern French, Occitan and related possibilities discussed.
What matters most for the master is not winning that dialect argument.
It is recognising the layer.
The month names appear to be later annotations. They therefore tell us that a later reader recognised or imposed a zodiacal-month relationship on these pages. That is historically useful.
They do not automatically identify the language of Voynichese.
A readable annotation can explain how a later reader understood a page without translating what the original scribe wrote.
Our dedicated Zodiac Pages article develops this distinction in detail.
Astrology and Medicine Were Historically Neighbours
Modern readers sometimes assume that “medical” and “astrological” are mutually exclusive categories.
That assumption does not fit the fifteenth-century intellectual world.
Medical miscellanies could contain calendars, zodiac diagrams, planetary material, regimen, bloodletting information, recipes, herbal material and practical advice within the same manuscript culture. Astrology could be used to organise timing, bodily correspondences or therapeutic decisions.
This means a manuscript containing both plant and zodiac imagery is historically plausible without requiring one of those visual regimes to be “fake”.
But historical compatibility is not specific identification.
A known medieval manuscript proves that certain combinations existed. It does not prove that Voynich uses those combinations for the same reason.
The Circular Pages Teach Us to Separate Shape from Function
Not every circle in a medieval manuscript is a star chart.
Circles are powerful information structures because they naturally represent recurrence, hierarchy, inclusion, sequence without a privileged beginning, concentric levels, directional relationships and cyclic time.
Medieval intellectual culture used circular diagrams for astronomy, calendars, cosmology, theology, medicine, logic, memory and other classificatory purposes.
Therefore the mere presence of a wheel does not tell us which of those functions is active.
The correct approach is relational:
- What occupies the centre?
- How many rings are present?
- What repeats in each ring?
- Where are labels placed?
- Does the page imply direction?
- Are there consistent counts?
- Do neighbouring diagrams preserve the same grammar?
- Does the text distribution change with the geometry?
A circle is not a meaning. It is a machine for arranging relationships.
Then Comes the Rosettes Foldout
There are Voynich pages that are strange because we cannot identify what they show.
The Rosettes foldout is strange because we can identify too many things it might show.
When unfolded, the large sheet presents nine major circular or rosette-like regions arranged approximately as a three-by-three field, connected by bands, paths, tubes, bridges or causeway-like forms. Around and between them sit walls, towers, cloud-like structures, patterned fields, star-like or flame-like motifs and numerous blocks of Voynichese text.
The central rosette is larger. The corner and edge rosettes differ from one another. The connections are not merely decorative filler. They create a network.
That network is the first thing we should trust.
Nine Things Connected Is Stronger Evidence Than One Castle-Like Shape
Individual Rosettes motifs have inspired geographical identifications for generations. Towers or battlements have been compared with real architectural traditions. Circular enclosures have been read as cities. Connecting bands have been read as roads, rivers, pipes or cosmic pathways.
Each may be worth testing.
Yet the most durable observation is simpler: the foldout encodes relationships among nine differentiated regions.
That statement survives whether the diagram is geographical, cosmological, anatomical, mnemonic, procedural or hybrid.
This is a recurring technique in strong Voynich work: move one level upward from the seductive identification to the structural fact that all serious identifications must explain.
Map, Cosmogram, Memory Palace or Process Diagram?
The Rosettes foldout can plausibly evoke several historical genres.
A map because regions connect spatially.
A cosmogram because circles, layered boundaries and celestial-looking motifs can represent ordered worlds.
A memory structure because differentiated places connected in a stable geometry can serve mnemonic navigation.
A process diagram because pathways can encode movement from state to state.
A composite because medieval diagrams did not always obey our modern disciplinary boundaries.
To choose among these, a proposal must explain more than visual mood. It must account for repeated motifs, text placement, directional relationships, the nine-region geometry and connections to other manuscript pages.
The Rosettes Problem Is Also a Reading-Order Problem
On an ordinary page, “start at the top left” is a reasonable default.
On the Rosettes foldout, that default may be meaningless.
Text appears around circles, between them and in differently oriented blocks. Some writing may need the sheet rotated for comfortable reading. A modern transcription must choose an order even when the manuscript does not provide a single obvious one.
This matters because sequence is part of interpretation. If we linearise a spatial document, we can accidentally manufacture relationships that were never intended.
The Rosettes foldout therefore teaches a principle that will become central in Part III:
Transcription is not neutral when the original object is spatial.
See the dedicated Rosettes Foldout article for the full geometry-first treatment.
What the Circular World Gives Us
By the time we leave the zodiac and Rosettes pages, the visual evidence has taught us several things without translating a word.
- The manuscript deliberately uses radial and networked information structures.
- Repeated human figures can function as units within a larger diagram rather than as narrative characters.
- Stars can be attached to figures, used as markers or incorporated into larger cosmological-looking fields.
- Some circular pages have externally recognisable anchors such as zodiac signs.
- Later readers added readable month names, proving the manuscript continued to be interpreted after its original production.
- The large foldout encodes a complex spatial relationship that cannot safely be reduced to ordinary page order.
- Visual identification becomes strongest when the recognised object predicts the architecture around it.
And then the manuscript changes again.
The circles give way to bodies.
Quire 13: When the Human Body Becomes Part of the Diagram
The transition into Quire 13 is one of the moments when Voynich seems to change genres in front of us.
Plants recede. Zodiac emblems recede. Large circular diagrams recede. In their place come human figures, pools, vessels, tubes, channels, openings, bulbous structures, flowing forms and dense blocks of text.
The figures are usually small. Many appear feminine. Some stand waist-deep or chest-deep in coloured pools. Some hold objects. Some appear to emerge from or enter tube-like structures. Some occupy margins as though the page itself were a plumbing diagram populated by people.
It is no surprise that generations of researchers have called this material biological or balneological.
Both names are useful.
Neither is a translation.
Why “Nymph” Is a Convenient Word, Not an Identification
Voynich researchers often call these small human figures nymphs. The term is compact, memorable and useful in transcription systems.
It also carries baggage.
A “nymph” in mythology is not simply a neutral word for a small drawn woman. If we forget that the term is a research convenience, we can unconsciously give the figures a mythological identity the manuscript itself has not established.
The safer description is more boring and therefore more useful: small human figures, many apparently feminine, placed within or beside diagrammatic structures.
Boring descriptions are underrated. They keep the evidence clean long enough for interpretation to earn its way in.
The Bodies Are Not Merely Bathers
On some folios, the visual scene genuinely resembles bathing. A green or blue pool contains several figures. Water-like colour surrounds their lower bodies. Medieval Europe had therapeutic bathing traditions and manuscript literature about mineral springs and baths.
That comparison is historically serious.
But many Quire 13 figures are not simply sitting in a bath.
They interact with tubes. They occupy compartments. They stand in structures that resemble containers. They may be paired with labels. On some pages the tube systems connect upper and lower regions. On f76v, a stream-like element crosses the binding into the related part of the bifolium at f83r—a physical reminder that some visual relationships belong to the sheet rather than to our modern single-page view.
This changes the question.
Instead of asking only “What are these women doing?”, ask:
- Why are particular figures placed at particular junctions?
- Why are some inside pools and others inside or beside tubes?
- Why do certain figures have individual labels?
- Why do some structures connect across the physical sheet?
- Why do colour, containment and flow recur?
- Are the figures people in a scene, units in a system, personifications, anatomical proxies, mnemonic actors—or several of these at once?
The moment we ask those questions, the figures stop being illustrations added to text and become components of a visual grammar.
The Quire Is Not Visually Uniform
Calling Quire 13 the “bathing section” can make it sound as though every page repeats the same scene.
It does not.
Some pages are crowded with pools and labelled figures. Some are dominated by long paragraphs with only marginal structures. Some contain isolated labels for tubes or tubs. One page, f76r, is text-only and contains a vertical sequence of stand-alone characters in the margin. Its ornate first character has led to the reasonable suggestion that it may once have functioned as a section opening.
This matters because the visual category and the textual category are not identical.
A “biological” quire can contain a page with no biological picture at all.
That should make us cautious about assuming the illustrations exhaust the subject of the text.
Labels Change the Evidentiary Game
Quire 13 does something especially valuable: it places short Voynichese strings near repeated visual units.
In René Zandbergen’s locus classification, the biological/balneological material includes dozens of labels associated with human figures and dozens more associated with tubs or tubes. His current transliteration summary counts 63 “nymph” labels and 47 labels for “tubes and tubs”. The exact category boundaries are editorial decisions, but the underlying phenomenon is visible on the manuscript: short pieces of text repeatedly sit beside differentiated visual elements.
That is precious because a label has fewer plausible discourse functions than a full paragraph.
A paragraph might describe, narrate, instruct, list, explain, calculate or encode something else. A short string adjacent to one figure is more likely to identify, classify, number, qualify or address that figure or its position—although even that remains a family of possibilities rather than a translation.
This is why label populations deserve their own analysis rather than being mixed indiscriminately with running text. Part III will return to the textual consequences. For now, the visual consequence is enough:
The manuscript repeatedly creates addressable visual units.
A figure can be one unit. A tube can be another. A star can be another. A plant fragment can be another.
Voynich is not merely drawing scenes. It often seems to be indexing parts of scenes.
Historical Bathing Is a Real Comparator
We should not dismiss the balneological comparison merely because it has sometimes been overstated.
Medieval Europe had a substantial tradition of therapeutic bathing. Pietro da Eboli’s De balneis Puteolanis, composed around the end of the twelfth century, described the therapeutic virtues of the thermal springs around Pozzuoli. Surviving illustrated copies show people immersed in baths, sometimes in circular or polygonal pools, participating in a medical world in which mineral waters were understood to have different effects.
Even closer to Voynich’s material date, the physician Ugolino da Montecatini—Ugolino Caccini—produced a Tractatus de balneis in the early fifteenth century, describing therapeutic waters, indications, methods and timing.
So there is nothing historically absurd about an early-fifteenth-century manuscript linking bodies, water and medicine.
That is the right conclusion.
The wrong conclusion is:
Medieval bath manuscripts existed; therefore Quire 13 is a copy of a bath manual.
The Voynich structures are often more interconnected, more tubular, more diagrammatic and less architecturally ordinary than the bathing scenes in well-known balneological manuscripts. A comparator demonstrates historical possibility. It does not erase difference.
That difference may be where the useful information lives.
Body, Water, Container, Flow
If we temporarily stop naming the subject and instead inventory recurring visual relations, Quire 13 becomes clearer.
- Body — repeated human figures, often differentiated by pose, position or object held.
- Water-like field — green or blue coloured zones that contain figures.
- Container — bounded pools, basins, vessel-like cavities or enclosed spaces.
- Connector — tubes, pipes, channels, arches or neck-like forms joining regions.
- Flow — implied movement through or between connected structures.
- Label — short text attached to some figures or structures.
- Paragraph — longer text occupying the remaining page field.
Those seven elements recur strongly enough that any serious interpretation should account for them together.
A pure “women bathing” reading explains body + water + container rather naturally but may struggle with elaborate connector systems.
A pure anatomical reading may explain body + flow + tube but must explain why the bodies are often external figures rather than anatomical cutaways and why pool-like compartments recur.
A reproductive or gynaecological reading can find suggestive body forms and enclosed spaces but must survive pages where the geometry seems more hydraulic than anatomical.
An alchemical reading may accommodate vessels and transformations but needs independent evidence that the diagrammatic actors correspond to substances or processes.
A cosmological reading can treat bodies as personifications and flows as cosmic relationships, but it must connect convincingly to the more explicitly celestial pages.
A mnemonic reading can treat figures as memorable anchors in a spatial system, but then the arrangement should have a repeatable logic that supports recall.
Each explanation captures something.
That may mean one is right and the others are superficial.
Or it may mean our modern categories slice the manuscript differently from its makers.
The Fifteenth Century Did Not Respect Our Departmental Boundaries
A modern university places botany, medicine, astronomy, astrology, hydrology, pharmacology and cosmology into sharply different intellectual boxes—some scientific, some historical, some no longer accepted as science.
A fifteenth-century compiler did not inherit those boxes.
Plants could be discussed because they altered the body. The body could be treated according to seasons. Astrological timing could influence medical decisions. Waters could possess therapeutic qualities. Minerals, stars, humours, temperatures and bodily states could participate in one explanatory world.
This does not give Voynich permission to mean anything whatsoever.
It gives us permission to stop insisting that a page must choose exactly one modern department.
Could These Be Anatomical Diagrams?
The tube systems have repeatedly encouraged anatomical readings: vessels, ducts, intestines, reproductive tracts, organs or physiological flows.
There are good reasons for the temptation. Organic tubes connect cavities. Bodies occupy the same visual system. Some forms look almost bodily.
But resemblance must survive anatomy.
If a structure is identified as a specific organ, its position, connection, repeated morphology and relationship to neighbouring structures should behave like that organ often enough to distinguish the claim from visual pareidolia. One successful “this looks like a uterus” comparison is not a reproductive-system model.
Likewise, an anatomical explanation must account for the page’s labelled figures and its pools, not only the one tube that resembles a duct.
The standard should be system-level coherence.
Could These Be Reproductive or Gynaecological?
The prominence of apparently female figures has invited interpretations involving pregnancy, menstruation, fertility, conception, childbirth and women’s medicine.
Again, the historical context does not make those subjects impossible. Medieval medical manuscripts discussed generation, women’s health, fertility and childbirth.
But a female body is not itself a diagnosis of subject matter.
The figures may be feminine because the represented domain concerns women. They may also be personifications, conventional diagrammatic figures, celestial or calendrical units, bathers, mnemonic agents or a visual population whose gender had symbolic rather than clinical significance.
Here, too, the question is not whether a reproductive interpretation can be made to fit one image. It is whether it predicts the organisation of many pages.
Could These Be Hydraulic?
Sometimes the simplest visual description is surprisingly productive.
Whatever the tubes signify, the manuscript is interested in connection.
One compartment joins another. A figure sits at a junction. A narrow neck leads into a larger basin. Colour differentiates one field from another. A cross-fold flow joins physically separated regions.
That does not mean “this is a water-engineering manual”. It means the visual grammar contains something flow-like.
This may be literal liquid flow, bodily flow, conceptual transition, categorical inheritance, sequence, influence or another relational mechanism.
We should not translate the connector before we have established what it connects.
The Human Figures Already Appeared in the Zodiac World
Quire 13 does not introduce the idea of repeated human units from nowhere.
The zodiac pages also organise numerous small human figures around central signs, often with stars and labels. The visual environment differs dramatically, but a recurring strategy survives: a human figure becomes a repeatable labelled unit inside a larger information structure.
This cross-regime recurrence is worth more than the superficial question “Are these the same women?”
The more powerful question is:
Does the manuscript use human figures as a general-purpose device for representing indexed entities?
If so, the figure may be less like a portrait and more like an icon.
That possibility would help explain why similar little bodies can live comfortably in very different visual worlds.
The Body Can Be a Unit Without Being a Person
This is easy to understand if we think about modern diagrams.
A stick figure on an airport sign is not a portrait of a traveller. A person icon in a population infographic is not one named citizen. A human outline in a medical diagram can stand for a class of patients. A silhouette in a flowchart can mean “user”.
Medieval visual systems were different from modern interfaces, but the general cognitive move is ancient: use a recognisable figure as a stable token.
Voynich’s repeated figures may therefore occupy a spectrum:
- specific person;
- type of person;
- body state;
- celestial or calendrical unit;
- personification;
- mnemonic marker;
- diagrammatic token.
The surrounding architecture should help us decide which end of that spectrum is more plausible on each page.
The Strange Case of f79v: When One Image Can Become Anything
Folio 79v contains one of those small Voynich moments that perfectly demonstrates the danger of visual recognition.
A figure appears in or near the mouth of a fish-like creature. Around it are other animals and watery structures.
The scene has been compared with a mermaid, a female Jonah-and-the-whale motif, Melusine and other mythic or religious imagery. Nearby animals have also attracted specific identifications.
Every comparison changes the story instantly.
Mermaid → mythology.
Jonah → biblical narrative.
Melusine → European legend and dynastic symbolism.
The image has not changed. Only the label in our head has.
This is why one unusual motif cannot carry an entire interpretation. A useful identification should explain neighbouring structures, text behaviour and recurrence elsewhere.
The Quire 13 Result That Survives the Competing Stories
After removing the claims that run ahead of evidence, something substantial remains.
- Human figures are a major repeated visual unit.
- Many figures are associated with coloured pools, bounded spaces or tubular structures.
- Connectors create relationships among compartments.
- Some relationships extend across the physical bifolium.
- Short labels attach to individual figures and structures.
- Long prose-like paragraphs coexist with those local labels.
- The same broad manuscript also uses labelled human figures in zodiacal diagrams.
- The section’s text belongs predominantly to Currier B in conventional classification.
- In Lisa Fagin Davis’s five-hand model, these pages are associated principally with her Hand 2; that attribution remains a palaeographic model rather than a semantic identification.
Those are strong constraints.
A successful explanation does not need to begin by saying “this is anatomy” or “this is bathing”. It can begin by explaining why body + containment + connection + label + paragraph recur together.
That is a much harder question.
It is also a much better one.
For the dedicated deep dive, see The Human Figures, Pools and Tubes.
From Bodies to Ingredients?
After the dense human world of Quire 13, the manuscript changes scale again.
The little bodies largely disappear.
Plant material returns—but now often in pieces.
And beside those pieces stand some of the strangest containers in the book.
The Vessels and Plant Fragments: When Voynich Stops Showing the Whole Thing
One of the most important transitions in the manuscript is easy to miss because the ingredients look familiar.
Plants return.
But they return differently.
On folios commonly called pharmaceutical, large specimen-like plants give way to rows of smaller roots, leaves and composite plant fragments. Beside them stand elaborate containers—tall, narrow, bulbous, tiered, decorated, sometimes almost architectural. Short labels attach to many fragments and vessels. Paragraphs occupy the spaces between rows.
The page no longer says visually: Here is one big thing. Look at it.
It says something closer to: Here are many units. Keep them separate.
That change in representational scale may turn out to be more important than the modern word “pharmaceutical”.
Folio 88r Is an Information Architecture Lesson
Take f88r as an example. The page contains three ornate container-like drawings and twelve miniature plant fragments arranged in four rows. The transcription map distinguishes sixteen lines of paragraph text, three container labels and twelve plant-fragment labels.
We do not need to know a single label’s meaning to see the organisation.
- objects are separated spatially;
- many objects receive short textual addresses;
- larger blocks of writing sit between object rows;
- container-like forms occupy a recurring visual role;
- plant-like forms are represented as fragments rather than complete specimens.
That is a document built for differentiation.
The page could be differentiating ingredients, preparations, categories, storage classes, names, properties or something else. The semantic label remains open. The document function is already more constrained.
Why the Containers Look So Persuasive
To modern eyes, the vessel forms invite the apothecary immediately.
Medieval and Renaissance pharmacies did use ceramic, glass and metal containers for medicines, herbs, syrups, ointments and compounds. Historical manuscripts also connect materia medica with recipes, ingredients, measures, substitutions and preparations. So “pharmaceutical” is not an absurd analogy.
But an ornate vessel is still not a prescription.
A container can signify at least four different things:
- storage — where something is kept;
- preparation — where something is mixed, heated, infused or processed;
- classification — a visual category marker rather than a literal vessel;
- symbolic container — a bounded unit representing a concept, state or quantity.
To choose among them, we need recurrence.
If vessel shape maps to a stable semantic category, similar shapes should predict similar labels, neighbouring object types, text behaviour or other measurable regularities. If every vessel can mean something different whenever a theory needs it to, vessel symbolism has ceased to explain anything.
The Plant Fragments Are More Interesting Than Their Species
The small plant drawings create a second temptation: identify each fragment botanically and then reconstruct a pharmacy inventory.
That may eventually be possible for some items. But the stronger first question is structural.
Why fragment the plant at all?
A whole plant and a root fragment do not merely differ in size. They imply different operations of attention.
A whole plant asks the viewer to recognise an organism.
A fragment asks the viewer to distinguish a part.
That distinction matters enormously in medicine, cooking, trade, agriculture and craft. Humans often encounter a living organism as a whole and use it as a part: bark, root, leaf, bulb, seed, resin, flower, fruit.
The Voynich pages may therefore be shifting from identity of source toward identity of usable component.
That interpretation is attractive.
We should still hold it at the correct level: plausible document logic, not decoded purpose.
Internal Reuse Is More Powerful Than External Resemblance
Several miniature fragments have been compared with larger plants elsewhere in the Voynich Manuscript itself.
On f89r1, for example, fragments have been noted as visually similar to larger plants on folios such as f57r, f30r and f33r. Other pages contain fragments compared with plants on f10v, f24r, f48r, f48v, f90r2, f94v and elsewhere.
Not every proposed match is equally convincing. Some are merely suggestive. Some may be the inevitable result of simple botanical shapes recurring.
But internal comparison has one major advantage over an external species hunt.
The manuscript itself becomes the reference system.
If a miniature fragment repeatedly matches a distinctive portion of a large plant and its label relates systematically to that plant page’s vocabulary, then the relationship becomes testable without first knowing the modern species name.
Internal identity can be established before external identity.
That is a powerful research strategy because it asks the Voynich Manuscript to define its own categories before we impose ours.
What Would a Strong Whole-Plant ↔ Fragment Match Give Us?
Suppose future work establishes, with high confidence, that one miniature root is a deliberate reuse of a root from a specific herbal folio.
That still would not translate the label.
But it would create a valuable internal equivalence.
- Whole-image page X and fragment unit Y belong to the same internal referent.
- The text attached to X can be compared with the label attached to Y.
- Repeated token components can be tested for shared distribution.
- Other fragments with similar visual relationships can be checked.
- The role of vessels on Y’s page can be tested against the inferred referent.
This is how pictures can constrain text without pretending to read it.
The Pharmaceutical Pages Are Also a Label Laboratory
Because so many discrete objects have adjacent short strings, these pages help us ask whether labels behave differently from ordinary prose.
If labels name objects, we might expect certain properties:
- shorter average length;
- different frequency profile;
- more rare forms if many objects have unique names;
- repetition when the same object class recurs;
- different positional constraints because labels do not need sentence grammar;
- possible overlap with keywords in related paragraphs.
Some of those questions can be—and have been—approached statistically.
But there is a prior caution: what we call a “label” is itself partly a layout judgment. A short isolated string near an object is probably functioning differently from a paragraph, but we should not assume every isolated string names the nearest thing.
Spatial attachment is evidence.
It is not a dictionary.
The dedicated articles The Vessels and Plant Fragments and Labels Versus Running Text take these problems much further.
Then the Pictures Almost Disappear
The final large regime presents a wonderful reversal.
After a manuscript full of plants, circles, bodies, pools and containers, the images nearly vanish.
What remains in the margin is deceptively simple.
Stars.
Quire 20: The Starred Pages
Quire 20 contains the manuscript’s final major text regime. It probably once consisted of seven normal-sized bifolia. The central bifolium—folios 109 and 110 in the old foliation—is now missing. The surviving leaves run from f103 to f108 and then f111 to f116.
The pages contain long stretches of Voynichese divided into entries of varying length. Along the margins sit star-like markers, many with tails and many with coloured centres.
René Zandbergen’s current quire description counts 324 surviving marginal stars. Because the central bifolium is missing, the original total is unknown.
That number is less important than the behaviour.
The stars are repeated markers attached to a predominantly textual environment.
The visual world has become punctuation-like.
Why Researchers Call Them “Recipes”
The conventional name comes from resemblance in document structure. Historical medical manuscripts often contain short recipes or practical entries separated by paragraph marks, initials, stars or other marginal signs. A long sequence of compact entries therefore looks recipe-like.
That is a sensible analogy.
It is not a decipherment.
The entries could be recipes. They could also be observations, instructions, catalogue records, prayers, mnemonic units, formulae, cases, names with descriptions or another genre that produces repeated short textual units.
The name “recipe section” should therefore be heard as “the section whose layout resembles collections of short recipes”, not “the section whose text has been translated as recipes”.
A Crucial Correction: One Star Does Not Always Equal One Paragraph
This is exactly the kind of detail that protects us from an attractive oversimplification.
On some pages the alignment between stars and short paragraphs is close. On others it is not.
Folio 103r, for example, has nineteen stars in the left margin while the paragraph count is estimated at around eighteen. On f106v there are fourteen stars but roughly fifteen textual paragraphs. On f108v the disparity becomes much stronger: sixteen marginal stars accompany a text that may have only about eight paragraphs, with several late stars running beside one long textual block. On f111r and f111v, stars again fail to map one-for-one onto obvious paragraphs.
This matters.
If we decide in advance that every star is a bullet point, the manuscript itself contradicts us.
The stars are clearly organisational marks. Their exact unit of organisation may vary or may be obscured by paragraph segmentation, later damage, writing sequence or our transcription conventions.
A marker can organise text without functioning as a simple one-to-one bullet.
The Stars Have Their Own Visual Variation
The marginal marks are not perfectly identical.
Many have tails. Some have no tail. Some have red centres. Others have faded yellow centres. On certain pages red and yellow appear to alternate for long runs. Some stars differ in size or orientation.
It is tempting to treat every variation as a code.
Perhaps colour marks categories. Perhaps tail type marks hierarchy. Perhaps alternating colours simply help the eye track adjacent entries. Perhaps some variation is decorative or production-related.
The responsible question is the same one we used for vessels: does the variation predict something else?
- Do red-centred stars precede a distinct vocabulary?
- Do tailed and untailed stars differ by entry length?
- Does colour correlate with scribal hand?
- Does a change in star form align with a change in textual regime?
- Does the pattern survive across pages rather than only within one attractive sequence?
Until such relationships are demonstrated robustly, star variation is a feature to measure, not a semantic legend to invent.
When Pictures Disappear, Structure Does Not
This may be the most important lesson of Quire 20.
Earlier parts of the manuscript use images to break the page into units. A plant divides paragraphs. Zodiac rings create repeated sectors. Human figures and tubes create labelled components. Plant fragments occupy rows.
In Quire 20, most of that figurative machinery disappears.
But segmentation remains.
The manuscript still wants the reader to see units.
A star in the margin may therefore be doing, in a highly compressed way, a job previously done by a larger picture: this is an addressable item.
That interpretation remains functional rather than semantic, which is precisely why it is useful.
From Illustrated Object to Abstract Marker
Seen across the manuscript, there may be a broad progression in how information units are represented.
- large plant — one dominant pictured object;
- zodiac figure — repeated pictured units inside a ring;
- human figure — repeated labelled units inside a connected system;
- plant fragment — repeated labelled components arranged in rows;
- star marker — abstract repeated sign beside text.
We should be careful with the word “progression” because the present manuscript order may not perfectly preserve production or conceptual order.
Still, the representational spectrum is real.
The book knows how to represent an information unit as a detailed object, a human icon, a small fragment or a nearly abstract star.
That is a clue about the manuscript’s visual literacy.
Why Quire 20 Is Not “Just Text”
A page can be visually designed without containing a figurative illustration.
Quire 20 uses margins, stars, colour, paragraph starts, indentation and line length to create structure. The absence of plants or people does not mean the page has stopped communicating visually.
Typography is visual information.
Spacing is visual information.
A star in a margin is visual information.
This is the bridge from Part II into Part III. Once the pictures shrink into marks and the marks sit beside writing, the border between “image” and “text” becomes less secure than it first appeared.
What Survives from Vessels to Stars
The page families look radically different, but a common informational concern remains visible.
Segmentation.
The manuscript repeatedly distinguishes one thing from another.
- one plant from its surrounding prose;
- one zodiac unit from the next;
- one human figure from another;
- one tube or vessel from another;
- one plant fragment from another;
- one starred textual unit from another.
This does not tell us what the units mean.
It tells us that the maker cared about units.
And once a manuscript cares about units, we can ask whether its writing system cares about units in corresponding ways.
That is where the visual world begins to hand the problem to the text.
For deeper treatment, see The Starred Pages.
DON’T GO IN CIRCLES: The Visual Edition
We can now freeze several visual mistakes so the next reader does not have to rediscover them by spending months inside the wrong theory.
“I recognised the plant, so the nearby word must be its name.”
No. First establish that the plant identification is distinctive. Then ask whether the nearby string behaves like an object label elsewhere. A plausible picture match plus proximity is a hypothesis generator, not a dictionary.
“There are zodiac signs, so the whole book is astrology.”
No. Zodiac imagery gives a real semantic frame to those pages. It does not assign the same function to herbal, Quire 13, pharmaceutical-looking or starred-text regimes.
“The month names look Romance, so Voynichese is Romance.”
No. The readable month names belong to a distinct later hand. They are useful evidence about reception and annotation precisely because they should not be collapsed into the primary script.
“The Rosettes foldout looks like my city.”
Then require the city hypothesis to predict the nine-node geometry, the connections, repeated motifs, orientation and textual placement. A castle-like form plus a favourite geography is resemblance, not provenance.
“Women in green water means baths.”
Bathing is historically plausible. The network topology, labels and cross-sheet relationships still require explanation. Balneology is a serious comparator family, not a recovered caption.
“Those are jars, therefore those are medicines.”
They are vessel-like forms compatible with apothecary comparison. Exact contents, functions and labels remain undecoded.
“The last pages are recipes.”
The safe result is stronger because it is narrower: Quire 20 contains a deliberately segmented text regime with repeated marginal star markers. Recipe semantics remain unproved.
“Everything connects, so the grand pipeline is solved.”
No. We examined the attractive plant → selected part → preparation → body → astrology → recipe story as a family of possibilities. Several transitions remain interesting; the universal pipeline did not generalise strongly enough to become the manuscript’s demonstrated architecture. Present binding order is not guaranteed conceptual order, and the cross-regime evidence does not yet force one sequential workflow.
What Part II Leaves Standing
- The manuscript has multiple genuine visual regimes.
- The modern names for those regimes are useful descriptions, not translations.
- Page architecture often survives interpretation better than exact object naming.
- Zodiac signs are strong external iconographic anchors while their Voynichese labels remain unread.
- The Rosettes foldout encodes relational geometry even though its external referent remains unsettled.
- Quire 13 repeatedly combines body, containment, connection, labels and prose.
- The vessel pages change representation from whole plants toward smaller, addressable components.
- Quire 20 preserves explicit entry-like segmentation even when figurative illustration largely disappears.
- Internal visual recurrence can produce useful constraints before external species, city or subject identification succeeds.
- No single visual theory currently earns the right to explain all of these regimes at once.
The pictures have not translated the manuscript. They have done something better: they have narrowed the questions.
From What We See to What the Writing Does
Now the difficult part begins. It would be convenient if each visible sign were a letter, each space a word boundary, each token a word and each repeated ending a suffix. Convenience is not evidence. Yet Voynichese is far from featureless. Its signs have habits. Its lines have edges. Its labels behave differently from prose. Some combinations are common, some rare, some positionally restricted.
Part III — What Voynichese Actually Does begins with the first textual rule: before we ask what a character means, we have to decide what we think a character is.
Part III — What Voynichese Actually Does
There is a moment in almost every Voynich investigation when somebody points at a glyph and asks the question that seems most natural in the world:
What letter is that?
It may be the wrong first question.
Before a mark can be a letter, we have to decide where the mark begins and ends. We have to distinguish one glyph from two touching glyphs. We have to decide whether a decorative stroke changes identity. We have to decide whether two slightly different forms are the same sign written differently or two different signs. We have to decide whether a tall form crossing another stroke is one character, a ligature, an abbreviation, a compound, a positional variant or something else.
That is why the deepest textual lesson in Voynich begins one step before language.
Before translation comes representation.
The manuscript gives us ink. Researchers turn the ink into glyphs. Then we turn glyphs into transcription symbols. Then software turns those symbols into data. Every statistical result, clustering result, frequency table and proposed translation stands downstream of those decisions.
This does not mean transcription is arbitrary. It means transcription is a model.
EVA Is a Brilliant Convenience, Not a Decipherment
The most familiar modern transcription convention is the Extensible Voynich Alphabet, usually shortened to EVA. It gives readable Latin-keyboard labels to recurring Voynich glyph shapes. Researchers can therefore write forms such as daiin, chedy, qokeedy or ol without reproducing the original manuscript shapes every time.
This is enormously useful.
It is also easy to misunderstand.
An EVA d is not a claim that the glyph sounds like English d. EVA o is not the vowel /o/. EVA q is not automatically the Latin letter q. The letters are convenient names for shapes or shape classes.
If somebody reads EVA aloud as if it were ordinary Roman transliteration, they have already smuggled phonology into a system designed mainly to avoid doing exactly that.
Transcription ≠ transliteration ≠ translation.
Transcription records a representation of the marks. Transliteration normally maps one writing system into another while preserving sign values. Translation maps meaning. Voynich research often possesses the first and lacks the next two.
Our dedicated article EVA, Transcription and the Segmentation Problem exists largely to prevent those layers from collapsing.
How Many Glyphs Are There?
You might expect this to be one of the easiest questions in the manuscript.
It is not.
A small set of common shapes accounts for much of the writing, but rarer forms, ligature-like combinations, bench forms, tall gallows, unusual line-initial embellishments and ambiguous joins make the inventory model-dependent. A researcher who treats a complex form as one glyph obtains one alphabet. A researcher who decomposes the same form into two or three components obtains another.
This has statistical consequences. Suppose what one transcription treats as “character X” is actually a compound AB. Character frequencies change. Conditional probabilities change. Entropy estimates change. Word lengths change. Repetition changes. Candidate cipher mechanisms change.
A 2022 PLOS ONE study by Jan Matlach, Barbora Janečková and Dominik Dostál explored precisely this representational problem, arguing for a model in which several familiar Voynich forms could be treated as compounds or ligatures and proposing a smaller underlying set of elements. Their model is interesting because it shows how much can change when segmentation changes. It is not a settled decipherment, and its steganographic interpretation remains one hypothesis among several. Read the study.
The broader lesson survives regardless of whether that particular model is ultimately right:
A surprising statistic can sometimes be telling us about our alphabet model rather than about the hidden language.
The Gallows: Tall Signs With Tall Expectations
Among the most distinctive Voynich glyphs are the tall forms researchers call gallows. In EVA notation they are commonly represented by forms such as k, t, f and p. Some appear as simple tall glyphs. Others combine with bench-like structures or more elaborate pedestal forms.
They are visually dramatic, which has encouraged dramatic meanings.
Capital letters. Paragraph markers. Numbers. Operators. Abbreviation signs. Cipher controls. Astrological values. Structural modifiers.
Any of those might turn out to capture part of their function.
What survives before meaning is positional behaviour.
Gallows forms are not distributed as if every position in the text were equally hospitable. Some are especially noticeable near beginnings of paragraphs or lines. Complex gallows arrangements interact with other common glyph families. Their height also lets them occupy vertical space differently from short glyphs, creating a visual hierarchy visible even before transcription.
This makes them excellent evidence that where a sign occurs matters.
It does not make them a proven semantic operator.
The full character-architecture problem is developed in The Gallows Characters.
Spaces Are Visible. “Words” Are an Interpretation.
The Voynich script contains visible gaps. Most transcriptions use those gaps to divide the text into tokens, and those tokens are routinely called words.
This is sensible shorthand. It is not a proven linguistic fact.
A visual gap can separate ordinary words. It can also separate abbreviational groups, cipher units, syllabic chunks, rhythmic units, formula components or writing segments whose function does not match a modern lexical word.
Conversely, what looks like one continuous Voynich token could contain more than one underlying linguistic unit if the script uses ligatures, abbreviations or encoded compounds.
This means every analysis of “word frequency” inherits a segmentation assumption.
That does not make word-level analysis useless. Quite the opposite. Voynich tokens exhibit rich regularities. But the safest wording is:
The manuscript has stable space-delimited token structure. Whether those tokens correspond one-for-one to linguistic words remains open.
See Spaces and Word Boundaries.
The q-Series: One of the Most Tempting Prefixes in the World
EVA q is peculiar.
It strongly favours token-initial position and is very often followed by EVA o, producing the familiar qo- beginning. Forms such as qokedy, qokeedy and relatives make the pattern visually obvious even to a newcomer.
That behaviour naturally invites linguistic analogies. Perhaps qo- is a prefix. Perhaps it marks a grammatical class. Perhaps it is an article plus stem. Perhaps it is a cipher control. Perhaps q and o together form one functional unit. Perhaps q modifies whatever follows.
The distribution strongly supports the idea that q-series forms participate in a constrained positional system.
It does not tell us which semantic story is correct.
This is a good example of how far Voynich research can advance without translation. We can say with confidence that q does not behave like a freely distributed ordinary sign. We can map its neighbours. We can compare its prevalence by section, line position, Currier regime and scribal hand. We can test whether proposed mechanisms reproduce the behaviour.
We should stop one step before “q means X” until the rest of the system forces X.
See The Q-Series: Why qo- Lives at the Beginning of So Many Voynich Words.
The Line Is Not Just a Place Where Words Happen to Fit
This is one of the strongest lessons to carry into any future decipherment.
Voynichese behaves differently near the beginnings and endings of written lines.
Certain glyphs and token forms show positional preferences. Some forms are more frequent near line starts. Others are associated with line endings. Paragraph starts can show additional effects. These patterns are visible across enough text that the physical line itself cannot safely be treated as meaningless wrapping imposed after composition.
In ordinary modern prose, a word processor may wrap a sentence wherever the right margin happens to fall. The linguistic system does not usually care whether the word “therefore” lands at the beginning of a display line.
Voynich makes that assumption unsafe.
The line may interact with composition, encoding, abbreviation, formula construction, scribal procedure or discourse structure.
This is why a proposed plaintext that ignores line boundaries can look locally convincing and still fail manuscript behaviour.
If the visible system cares about line position, a complete explanation must tell us why.
The dedicated treatment is The Line as a Unit.
Paragraph Beginnings Are Not Ordinary Either
Paragraphs frequently begin with visually or statistically distinctive material. Tall gallows forms are especially noticeable near some paragraph openings. First-line behaviour can differ from later lines. Tokens at paragraph starts can have distributions that are not simply random samples from the rest of the paragraph.
This can be generated by many mechanisms.
- A paragraph marker or enlarged initial.
- A formula used to open records.
- A scribal convention.
- A cipher state reset.
- A grammatical construction.
- A structural token generated by layout rules.
- A decorative convention correlated with, but not semantic in, text structure.
The positional effect is evidence. The list of possible causes is interpretation.
Labels Are Voynichese, But They Are Not Ordinary Prose
We met labels visually in Part II: short strings beside stars, human figures, plant fragments, vessels and diagram components.
Textually, they create one of the best natural experiments in the codex.
The same writing system is being used in a different document position. If labels name, classify or index objects, they may not require the same grammar as running prose. Even without reading them, we can compare their length, token inventory, recurring forms and overlap with nearby paragraphs.
Labels do share substantial material with the wider Voynich lexicon. They are not an entirely separate alphabet pasted onto pictures. At the same time, their distributions and practical context differ enough that mixing all labels indiscriminately into prose analysis can erase useful structure.
The key rule is:
Same script does not imply same discourse function.
See Labels Versus Running Text.
Circular Text Reminds Us That Reading Order Can Be Invented by the Transcriber
A normal line offers an obvious sequence. A ring does not.
When researchers transcribe writing that circles a diagram, they have to choose a starting point and direction. When text is scattered around a foldout, they have to choose which block comes first. Those choices are necessary for a linear digital corpus.
They are not automatically the historical reader’s sequence.
This matters for n-gram analysis, local co-occurrence and any model that assumes neighbouring transcribed tokens were read in that order. Spatial text may encode relations by position rather than sequence.
A future computational treatment of Voynich should therefore preserve more than characters. It should preserve geometry: coordinates, orientation, label attachment, line identity, paragraph identity, folio, bifolium and diagram region.
What the Text Already Tells Us Without Meaning
By now we can make a surprisingly substantial list.
- The writing uses a limited, recurring inventory of visual forms.
- Exactly how that inventory should be segmented remains partly model-dependent.
- Space-delimited tokens have stable internal restrictions.
- Some signs strongly prefer particular positions within tokens.
- EVA q strongly favours token beginnings and typically associates with o.
- Gallows characters have distinctive positional and compositional behaviour.
- Line beginnings and endings are distributionally different.
- Paragraph beginnings can be distributionally special.
- Labels and running text overlap yet should not be assumed to have identical functional grammar.
- Spatial diagrams complicate any assumption of one obvious reading order.
- Transcription choices can materially affect downstream statistics.
None of those statements says what daiin means.
All of them constrain what a successful account of daiin can be.
DON’T GO IN CIRCLES: The Transcription Edition
“EVA d must sound like d.”
No. EVA is primarily a transcription convention. Sound values remain unknown.
“There is a space, therefore there is a word.”
A stable token boundary is real. Its linguistic interpretation remains open.
“The gallows are clearly capitals.”
Their position can make a capital-like analogy attractive. Other mechanisms can produce the same positional concentration. The analogy is not an identification.
“qo- is obviously a prefix.”
It behaves prefix-like at the visible token level. Its underlying mechanism and meaning remain unproved.
“A line-start pattern tells us what the opening glyph means.”
No. Positional behaviour establishes a role constraint, not a semantic value.
“The transcription file is the manuscript.”
No. It is a machine-readable model of selected manuscript features. The physical folio always has the right to overrule the data file.
Part III Reading Map
- What the Writing Does Before We Know What It Says
- EVA, Transcription and the Segmentation Problem
- The Gallows Characters
- Spaces and Word Boundaries
- The Line as a Unit
- Labels Versus Running Text
- The Q-Series
The Next Problem: There May Be More Than One Voynichese
So far we have spoken about “the script” as though one consistent statistical machine produced every page.
It did not behave that simply.
Vocabulary shifts. Character preferences shift. Pages cluster. Scribes—or at least hand styles—shift. Labels and prose diverge. Currier A and Currier B divide large parts of the manuscript into distinct textual populations.
The next question is therefore not “What language is Voynichese?”
It is:
What hidden variable makes one region of Voynichese behave differently from another?
Part IV — The Text Is Not One Uniform Thing begins there.
Part IV — The Text Is Not One Uniform Thing
Imagine discovering an unknown book in which every page uses the same unfamiliar script.
The natural assumption is that all the pages belong to one textual system.
Voynich punishes that assumption too.
The script is visually coherent enough that the manuscript feels unified. Yet when the writing is counted instead of merely looked at, different regions prefer different forms. Some token families flourish in one group and recede in another. Certain character combinations change frequency. Labels form another population. Page families cluster. Handwriting varies. Even a reader who cannot translate one token can detect provinces inside the text.
This gives us one of the most important surviving facts about Voynichese:
The manuscript is textually structured at more than one scale.
The hard part is deciding what creates those scales.
Currier A and Currier B: A Distinction That Refused to Go Away
In the twentieth century, Prescott Currier observed that large parts of the manuscript could be divided into two broad textual populations based on recurring statistical and orthographic differences. These became known as Currier A and Currier B.
The terminology can be misleading because they are sometimes called “languages”. That was convenient shorthand. It should not be heard as proof that two different spoken languages lie underneath the script.
The durable result is distributional:
- many pages cluster naturally into A-like or B-like textual behaviour;
- certain common forms have markedly different frequencies across the two groups;
- the distinction correlates partly with manuscript regions and page types;
- the two regimes share enough structure to remain recognisably Voynichese;
- not every page fits a clean binary without complication.
This is stronger than saying “Currier thought the manuscript had two languages.” The distinction has remained productive because later researchers can rediscover related clustering using different statistical tools.
Something changes.
The name of that something remains open.
See Currier A and B for the dedicated treatment.
What Could the Hidden Variable Be?
If A and B are real distributions but not necessarily two languages, what could generate them?
Several families remain plausible.
Different languages
The most literal interpretation is that the same script encodes two linguistic systems. That is possible, but it must explain why the systems share so much visible architecture and why their distribution interacts with page families and scribal variation as it does.
Different dialects or registers
One language can change across dialect, genre, technical register, time period or author. A medical recipe and an astronomical explanation can have very different vocabularies while remaining the same language.
Different scribes
Individual writers can prefer different spellings, abbreviations, glyph forms and formulae. If scribes copied from different exemplars or were trained differently, handwriting and textual statistics could shift together.
Different functions
A label, a paragraph, a list entry and a diagram caption may use the same underlying language differently. Currier-like clustering could partly reflect document function rather than speech community.
Different production stages
A long project can evolve. Conventions can drift. A scribe can learn. A cipher key can change. An abbreviation system can become more compressed. Later sections can inherit and modify earlier habits.
Different encoding modes
If the visible text is transformed, the transformation itself may have modes. The same underlying language could produce different surface distributions if parameters, alphabets, null rules, syllabification or abbreviation practices change.
These possibilities are not mutually exclusive.
A manuscript could have multiple scribes working in different sections, on different subject matter, at different times, using slightly different textual conventions. One observed partition can therefore be the shadow of several hidden variables acting together.
The Most Dangerous Shortcut: Currier A = Topic A, Currier B = Topic B
Because Currier groups correlate partly with visual sections, it is tempting to assign subject meanings directly.
A-text = plants.
B-text = bodies.
Or A = descriptive, B = procedural.
Those hypotheses can be tested, but the distribution does not license them automatically. Topic, scribe, period and document type can correlate. A statistical cluster tells us that pages differ; naming the cause requires independent evidence.
This is the same firewall we used for images:
Distributional cluster ≠ decoded topic.
The Scribe Problem: Five Hands Became a New Kind of Certainty
In 2020, palaeographer Lisa Fagin Davis published an influential digital palaeographic analysis arguing that the manuscript was written by five distinct scribes. Her work compared recurring glyph shapes and handwriting habits across the codex and gave researchers a practical hand classification that could be compared with Currier language and page type. The article, How Many Glyphs and How Many Scribes? Digital Paleography and the Voynich Manuscript, appeared in Manuscript Studies. Read the bibliographic record.
The five-hand model was useful because it supplied a new axis.
Now researchers could ask:
- Does Currier A correlate with one hand?
- Does Currier B correlate with another?
- Do certain image sections belong disproportionately to particular hands?
- Do vocabulary shifts persist when hand is controlled?
- Do labels and prose differ because different scribes produced them?
That is good science: a new classification generates new comparisons.
But useful classifications can harden into facts faster than their uncertainty disappears.
2026: The Five-Hand Model Was Challenged
In 2026, Torsten Timm published a direct critique titled One Hand, Five Labels: A Critical Examination of the Five-Scribe Hypothesis for the Voynich Manuscript. Timm argues that some of the diagnostic variation used to separate hands occurs much more continuously across the manuscript than discrete categories imply, and that several of the five labels align suspiciously closely with distinctions already present in Currier classification or text function. He proposes that continuous evolution within one hand may explain the variation more parsimoniously. Read the 2026 critique.
This does not mean “science proved there was one scribe”.
It means the five-scribe model is now an explicitly contested scholarly model rather than a number we should repeat without qualification.
The appropriate public conclusion in 2026 is:
Voynich contains meaningful handwriting variation. Whether that variation is best partitioned into five discrete scribes remains under debate.
That sentence is less exciting than “five scribes wrote Voynich”. It is more durable.
The dedicated article The Scribe Problem preserves both the original model and the later challenge.
Why Scribe Counts Matter Less Than Production Workflow
Suppose there were five scribes.
What would that actually tell us?
It would strongly suggest collaborative production or at least multiple writing episodes. But it would not by itself tell us whether those scribes composed, copied, translated, encrypted or merely transcribed material. A scribe can copy a language they do not understand. Five scribes can copy one author. One scribe can copy five sources.
Now suppose there were one principal hand whose forms changed over time.
That would suggest continuity, but it still would not prove one author, one subject or one language. A single person can spend years copying heterogeneous material.
So the deeper question is not:
How many people held the pen?
It is:
What production process best explains the interaction among hand variation, textual regimes, physical gatherings and visual programmes?
A headcount is only useful insofar as it helps answer that larger question.
Visual Section, Textual Regime, Scribe and Quire Are Different Axes
This may be the single most important organisational insight in Part IV.
Researchers often divide Voynich in several ways:
- by visual regime: herbal, zodiac, Quire 13, pharmaceutical-looking, starred;
- by textual regime: Currier A, Currier B, mixed or uncertain;
- by hand model: one or multiple scribal hands;
- by physical structure: bifolium and quire;
- by document function: labels, running paragraphs, circular text, star-marked entries;
- by local vocabulary: neighbouring pages that share unusually specific forms.
These axes overlap.
They do not collapse perfectly onto one another.
That is extremely important.
If the manuscript had one simple hidden division—say, five authors writing five topics in five separate quires—we would expect the partitions to align neatly. Instead, the intersections are messier.
Messiness is not evidence of chaos.
It may be evidence of a real production system with multiple interacting causes.
Think in Layers, Not Boxes
A useful analogy is a modern school.
Students can be grouped by class, year, subject, teacher, ability band, timetable and house. None of those classifications is fake. None is the one true partition of the school.
A student belongs simultaneously to several systems.
Voynich may be similar. A folio can simultaneously belong to:
- a physical quire;
- a visual page family;
- a Currier regime;
- a hand style;
- a local vocabulary neighbourhood;
- a discourse type such as labels or paragraphs;
- a production episode.
Searching for one master classification can therefore destroy information rather than simplify it.
We learnt to ask a better question:
Which partitions agree, which disagree, and what hidden cause would generate that pattern of agreement?
Local Vocabulary Creates Another Geography
Beyond the broad Currier divide, neighbouring pages can share unusually specific token families. Certain forms are concentrated in limited zones. Some words are common nearly everywhere; others behave like local residents.
This produces what we might call vocabulary geography.
A token can tell us something about where in the manuscript we are even when we cannot translate it.
That is a remarkable property.
In an ordinary book, topic-specific vocabulary does the same thing. Words such as “mitochondria” and “ribosome” make a biology chapter statistically different from a chapter on algebra. Names make one historical section different from another. Formulaic openings distinguish recipes from essays.
But the same phenomenon can also arise from different scribes, encoding modes or locally generated text.
Therefore local vocabulary proves localisation, not topic identity.
Our article Keywords Without Meanings develops this distinction.
Labels Form Their Own Province Too
Label text cuts across visual sections. We find short strings beside zodiac figures, stars, human figures, vessels, roots and diagram components.
That makes labels a functional category independent of subject.
If we discover a statistical feature shared by labels in herbal, zodiac and Quire 13 pages, that feature may belong to the grammar of labelling rather than the grammar of botany, astrology or medicine.
This is why mixing all text together can blur the mechanism.
A whole-corpus average can describe no actual region well.
Imagine averaging the language of a dictionary, a poem, a spreadsheet and a legal contract. The result may be mathematically correct and linguistically unrepresentative of every document.
Voynich research needs both scales: whole-manuscript regularities and local-regime differences.
The Manuscript May Be Dynamic Rather Than Discrete
One of the most interesting challenges raised by recent work is whether some categories we treat as boxes are actually points along continua.
A scribe’s handwriting can drift gradually.
Vocabulary can drift gradually.
A production convention can be learned, simplified or elaborated.
A cipher or notation system can change parameters over time.
If a continuous process is later divided into categories, the categories may still be useful summaries. The danger comes when we forget they were summaries and begin treating their boundaries as historical events.
“A becomes B on this page” is a much stronger claim than “these pages occupy different ends of a statistical continuum”.
This matters because a dynamic-text model and a multiple-language model make different predictions about transition pages, mixed folios and gradual changes.
What Would Resolve Currier A/B?
Not another label.
We need a mechanism.
A strong account should explain:
- which features distinguish A and B;
- why those features co-vary;
- why some pages are intermediate or difficult to classify;
- how the distinction interacts with visual regimes;
- how it interacts with hand variation;
- whether labels obey the same division;
- whether line and paragraph effects change across regimes;
- whether the same generative or linguistic rules can produce both.
Then the account should make a prediction.
If the difference is dialect, what new distribution follows?
If it is scribe, what happens on pages assigned to the same hand across different visual sections?
If it is production time, can transition be aligned with physical gatherings?
If it is cipher mode, what invariant survives underneath both surface forms?
That is how a classification becomes an explanation.
DON’T GO IN CIRCLES: The Classification Edition
“Currier A and B are two languages.”
They are two strong distributional regimes. Language identity is one possible cause, not the observed fact itself.
“Five scribes wrote the manuscript.”
Five hands remain an influential palaeographic model, but a 2026 critique argues that the variation may be continuous and compatible with fewer hands. The number is contested.
“If two pages have the same scribe, they have the same topic.”
A scribe can copy multiple subjects. Hand identity does not determine semantic identity.
“The visual sections are the textual sections.”
They correlate imperfectly. Illustration type and textual statistics are different axes.
“A statistical difference tells us who wrote the text.”
It tells us that something differs. Author identity requires palaeographic, historical or other independent evidence.
“One clean partition should explain the whole manuscript.”
That expectation may itself be wrong. Physical, visual, scribal, textual and functional partitions overlap without becoming identical.
What Part IV Leaves Standing
- Voynichese is not statistically uniform across the manuscript.
- Currier A/B remains one of the strongest broad textual distinctions.
- The cause of that distinction is not established.
- Handwriting variation is real, while the exact number of scribes remains debated.
- Labels form a cross-cutting functional text class.
- Local vocabulary creates smaller-scale textual geography.
- Visual sections, Currier groups, physical quires and hand models are overlapping but non-identical partitions.
- Some apparent categories may reflect continua rather than sharp historical boundaries.
- A successful theory must explain the intersections, not merely choose one partition and ignore the rest.
The manuscript is therefore neither one undifferentiated soup nor a set of neatly labelled boxes.
It is a structured field.
And inside that field, the individual “words” behave in ways stranger than the broad partitions alone can explain.
Part V — Words That Behave Strangely begins with a simple observation: if you look at enough Voynich tokens, they start to resemble families more than strangers.
Part V — Words That Behave Strangely
If Voynichese were merely unfamiliar, it would be easier.
We know what to do with unfamiliar languages. We collect text. We identify recurring signs. We compare contexts. We look for names, numbers, formulae, inflections, repeated phrases and bilingual anchors. Given enough material, ordinary linguistic structure usually begins to yield.
Voynich gives us enough material to see structure very clearly.
Then the structure behaves just strangely enough to stop being reassuring.
Many tokens look related to other tokens by tiny edits. One glyph disappears. Another appears. A beginning changes while the middle survives. A common ending expands. Nearby lines reuse similar shapes. A rare form may cluster on one folio. A token family can proliferate through a local region like variations on a theme.
At the same time, the system is highly constrained. Not every glyph can occur everywhere. Some combinations are common; others are extraordinarily rare. Given part of a token, the rest can sometimes be predicted unusually well compared with many ordinary alphabetic texts.
This is the point where two opposite intuitions collide.
It looks too organised to be nonsense.
It looks too organised to be ordinary language.
Both intuitions have inspired serious work.
Neither is a conclusion.
The First Strange Thing: Voynich “Words” Have Relatives Everywhere
Take a common Voynich token in EVA transcription. It is often not isolated in the vocabulary. Nearby are forms that differ by one small operation: one sign inserted, deleted, substituted or repeated.
Forms conventionally transcribed as chedy, sheedy, cheedy, qokedy, qokeedy and many others participate in visibly related families. The exact relationships depend on how glyphs are segmented and which transcription is used, but the family resemblance is one of the manuscript’s most persistent textual impressions.
Natural languages do this too.
English gives us walk, walks, walked, walking. Latin builds paradigms with stems and endings. Arabic root-and-pattern morphology generates families whose members share consonantal structure. Agglutinative languages can assemble long words from reusable pieces. Medieval scribes also used abbreviations that create families of visible forms.
So “Voynich has word families” is compatible with language.
But a generator can create families too.
Begin with a token. Copy it. Change one glyph. Copy the new form. Add a prefix-like element. Remove an ending. Repeat. A local-copying mechanism can produce a vocabulary in which every word seems to be cousin to another without encoding ordinary morphology.
A cipher can create families too if plaintext letters expand into multiple symbols, if homophones are selected under contextual rules, or if the encoding uses structured fillers and prefixes.
Therefore:
Word-family structure is real. Morphological meaning is one explanation of that structure, not the structure itself.
This is developed in the satellite article Word Families.
Near-Neighbours Are More Important Than Famous Words
Voynich discussions often become attached to individual famous forms: daiin, ol, chedy, qokedy. That is understandable. A named token feels like something we can hold.
But the more informative object may be the neighbourhood around the token.
Imagine a city street. One house tells us little about the planning law. A hundred neighbouring houses reveal setbacks, plot widths, common materials, permitted heights and repeated floor plans.
Voynich tokens work similarly. A token’s possible internal structure becomes clearer when we ask:
- Which one-edit neighbours exist?
- Which theoretically possible neighbours never occur?
- Which changes concentrate at the beginning?
- Which changes concentrate at the end?
- Does doubling a glyph preserve the rest of the frame?
- Do q-prefixed versions occur for only certain families?
- Does a family prefer Currier A or Currier B?
- Does it prefer labels, paragraphs or line boundaries?
- Does it spread through the manuscript or cluster locally?
A theory of morphology, abbreviation, encryption or generation should explain not merely which forms exist but the shape of the whole neighbourhood.
The Token Often Looks Like It Has Zones
Researchers have long noticed that Voynich glyphs do not combine with equal freedom. Some tend to occur toward beginnings, some toward middles, some toward endings. Certain sequences appear in stereotyped orders. Common endings recur across many token families.
This can make a Voynich token look as though it has something like:
beginning zone → core zone → ending zone
That arrangement naturally resembles morphology.
A prefix occurs first. A stem carries lexical identity. A suffix marks grammar.
But the same surface zoning can arise from other systems.
- An abbreviation convention can place category marks first and suspensions last.
- A cipher can encode word starts differently from middles.
- A syllabic system can restrict legal initial and final signs.
- A generated-token system can assemble components from ordered slots.
- A verbose cipher can expand plaintext symbols into multi-character groups with constrained positions.
- A writing system can combine base signs with modifiers that have fixed attachment order.
So the zones are evidence about combinatorics.
The labels “prefix”, “stem” and “suffix” remain hypotheses until a linguistic mechanism earns them.
Why Repetition Feels Wrong — and Why Feeling Is Not Enough
Voynich contains repetitions that often strike experienced readers as unusual. Similar tokens can appear close together. Near-duplicates can occur within a line or adjacent lines. Sequences can seem more self-similar than ordinary prose.
This inspired the idea that the text might have been generated locally by copying and mutating nearby tokens.
The intuition is powerful because the manuscript sometimes looks exactly like that.
Yet meaningful texts can repeat heavily too.
- Recipes repeat ingredients and formulae.
- Liturgical texts repeat invocations.
- Legal documents repeat fixed constructions.
- Medical lists repeat dosage or preparation language.
- Tables and catalogues repeat category markers.
- A verbose cipher can turn one repeated plaintext pattern into several visibly related ciphertext patterns.
Therefore the correct question is not “Does Voynich repeat more than the novel on my bedside table?”
The correct question is:
Which generative mechanisms reproduce the type, scale and location of Voynich repetition while also reproducing everything else the manuscript does?
Local Echo: Why Nearby Words Matter So Much
Some Voynich forms have an unusually local life.
A rare token may appear several times on one folio and scarcely anywhere else. A family may cluster across neighbouring pages. Nearby paragraphs can share vocabulary more strongly than distant regions do.
This could reflect ordinary topic.
A chapter about roses uses “rose” more than a chapter about planets. A medical recipe discussing one substance repeats that substance. Names cluster in the passages where their people appear.
But local echo can also be produced mechanically.
A scribe who generates new tokens from nearby tokens will create local families. A cipher with changing local parameters will create local distributions. A source manuscript copied section by section can preserve local vocabulary without the visible strings carrying straightforward lexical meaning.
The observed fact is localisation.
Topic is one candidate cause.
2013: Long-Range Structure Became Harder to Dismiss
In 2013 Marcelo Montemurro and Damián Zanette published an information-theoretic study titled Keywords and Co-Occurrence Patterns in the Voynich Manuscript. They examined how word-like tokens distribute across the manuscript and reported long-range organisation compatible with the kind of topical structure seen in meaningful texts. They also identified statistically significant co-occurrence networks and candidate “keywords”. Read the PLOS ONE study.
The study was important because it moved beyond simple frequency.
A random bag of symbols can accidentally produce common tokens. A stronger test asks whether particular tokens recur in coherent regions and whether their co-occurrence structure resembles organised text.
Montemurro and Zanette found evidence that the manuscript does have long-range structure.
That is a durable result to take seriously.
But the word “keyword” can mislead readers.
A statistical keyword is a token whose distribution helps distinguish one region of a corpus from another. It is not a translated noun.
If a token occurs disproportionately on herbal pages, we can say it helps classify herbal-page text. We cannot jump directly to “therefore it means plant”, “root”, “medicine” or a species name.
A token can tell us where it belongs before it tells us what it means.
That distinction is the core of Keywords Without Meanings.
Another 2013 Result: Voynich Is Not Well Described as Random Text
That same year, Diego Amancio and colleagues published Probing the Statistical Properties of Unknown Texts: Application to the Voynich Manuscript. Their framework compared first-order word statistics, text-network properties and intermittency measures across known-language corpora, shuffled controls and the Voynich text. They found the Voynich corpus broadly compatible with natural-language-like structure on several measures and unlike simple randomised controls. Read the PLOS ONE study.
This contributed to a growing body of evidence against the simplest “random gibberish” description.
But simple randomness was never the only non-plaintext alternative.
A deterministic generator is not random. A cipher is not random. A structured mnemonic notation is not random. A rule-governed pseudo-language is not random. A copy-and-mutate system is not random.
Therefore:
Non-random ≠ natural language identified.
The Natural-Language Lens Is Still Productive
Claire Bowern and Luke Lindemann’s 2021 review in the Annual Review of Linguistics brought together a large body of linguistic work on Voynichese. They argued that treating the text seriously as language remains productive and showed how questions about phonology, morphology, orthography and structure can be investigated statistically even without decipherment. Read The Linguistics of the Voynich Manuscript.
This is important because “we have not identified the language” does not mean linguistic analysis has failed.
A linguist can ask whether glyph sequences resemble phonotactics. They can ask whether token families behave like morphology. They can compare conditional entropy. They can model how many underlying units might be needed. They can test candidate languages and transformations. They can ask whether observed restrictions are compatible with known writing systems.
Those are legitimate linguistic questions.
The caution is not “do not use linguistics”.
It is “do not promote compatibility into identity”.
Language-like ≠ language identified.
Entropy: A Fancy Word for an Extremely Useful Question
Entropy sounds intimidating because it lives in several branches of science.
For text analysis, the practical intuition is simple.
How uncertain are we about what comes next?
If a writing system permits almost any character after any other character with roughly equal likelihood, the next symbol is hard to predict. If only a few characters are legal after a given context, the next symbol is more predictable.
Voynichese is unusually constrained at the character level under common transcription schemes. Certain sequences are highly favoured. Some signs almost never occur in particular positions. Character-level conditional entropy can therefore be strikingly low relative to many familiar alphabetic texts.
This has generated two opposite mistakes.
Mistake A: Low entropy proves meaningless generation
It does not. A transformed meaningful text can become more predictable at the surface. A writing system with strong positional spelling rules, verbose encoding, abbreviations or constrained glyph combinations can lower visible uncertainty.
Mistake B: Language-like entropy proves plaintext
It does not. A generator or cipher can deliberately or incidentally reproduce language-like statistical ranges.
Entropy is therefore a constraint on mechanism.
Any proposed mechanism should explain why the output is as predictable as it is.
Our dedicated Why Voynichese Is So Predictable develops the idea without pretending entropy is a decoder ring.
Predictability Can Come From Several Layers
Suppose the next visible glyph is unusually easy to predict.
Why?
- Language: the underlying language may have strong phonotactic or morphological constraints.
- Orthography: the writing system may represent sounds or syllables with constrained spellings.
- Abbreviation: scribal conventions may force particular signs into particular positions.
- Cipher: an encoding procedure may expand plaintext into predictable multi-symbol groups.
- Segmentation: transcription may split one functional sign into several visible characters, artificially increasing apparent predictability.
- Generation: a rule-based process may assemble tokens from a small set of ordered components.
- Layout: line-start or line-end constraints may change what symbols are available locally.
The same statistic can therefore be generated by different layers.
This is why a model that matches one entropy number has not solved Voynich.
It has passed one gate.
2025: A Cipher Demonstrated an Important Possibility
A particularly useful development arrived in 2025 when Michael A. Greshko published The Naibbe cipher: a substitution cipher that encrypts Latin and Italian as Voynich Manuscript-like ciphertext in Cryptologia. The paper constructs a verbose homophonic substitution cipher designed to be executable by hand with historically plausible materials and shows that encrypting Latin and Italian plaintexts can reproduce multiple Voynich-like statistical properties while retaining recoverable meaningful text underneath. Read the 2025 study.
This does not prove that the Voynich Manuscript uses the Naibbe cipher.
That would be exactly the kind of overreach this master is designed to prevent.
Its importance is methodological.
For years, researchers could point at a Voynich statistic and say: “Ordinary substitution ciphers do not look like this, therefore cipher is unlikely.” A constructed counterexample can change that inference. If a historically executable cipher family can generate several of the supposedly incompatible properties at once, then those properties no longer rule out all ciphering in the simple way claimed.
The cipher hypothesis survives more strongly.
The particular cipher remains unproved.
A model can prove possibility without proving identity.
Why Counterexamples Matter So Much
Suppose someone says:
No meaningful encrypted text could ever produce this combination of Voynich-like properties.
One counterexample is enough to defeat the word ever.
It does not tell us which historical mechanism produced MS 408. It changes the logical landscape by reopening a region someone thought had been eliminated.
This is a recurring feature of mature Voynich research.
Sometimes progress is not “we found the answer”.
Sometimes progress is “that argument can no longer exclude this family of answers”.
But Synthetic Similarity Has a Trap of Its Own
Once a generated or encrypted text resembles Voynich, the opposite overreach becomes tempting.
My algorithm produces Voynich-like statistics; therefore Voynich was made by my algorithm.
No.
Many mechanisms can converge on similar outputs.
Two different languages can have similar word-length distributions. Two cipher systems can have similar entropy. A copying generator and a morphological language can both produce edit-distance families. A model fitted to known statistics can reproduce those statistics because that is what it was built to do.
Mechanism identity requires more.
- Does it predict properties not used to construct the model?
- Does it reproduce Currier A/B differences?
- Does it reproduce line-position effects?
- Does it explain labels versus prose?
- Does it reproduce local vocabulary geography?
- Can it account for scribal or production variation?
- Does it yield meaningful, consistent output rather than only Voynich-like surface texture?
- Can another researcher reproduce the result from explicit rules?
Pattern match ≠ mechanism identity.
Zipf’s Law Is Not a Certificate of Language
One frequently discussed property of text is the rank-frequency relationship often associated with Zipf’s law: a few words are very common, many are moderately common and a long tail occurs rarely.
Voynich token frequencies show broadly language-like rank behaviour under common tokenisations.
That is interesting.
It is not unique to human language.
Zipf-like distributions can arise in many complex systems and can be produced by generative processes that do not encode ordinary semantic prose. A model should therefore not receive a “natural language” certificate merely for producing a long-tailed frequency curve.
The evidentiary value comes from combinations of properties, especially properties difficult to satisfy simultaneously.
Word Length Is Not One Simple Number Either
Voynich tokens are often described as having an unusual length distribution.
Before interpreting that distribution, remember Part III.
What is a character?
What is a word boundary?
If one transcription treats a ligature as two characters and another as one, token length changes. If a visible space separates syllabic or cipher units rather than lexical words, “word length” is measuring the wrong linguistic object. If the script is verbose, one plaintext letter may expand into several visible signs.
Word-length statistics remain useful because competing models must reproduce the visible distribution.
They should not be interpreted as if the visible unit were already known to be an ordinary alphabetic word.
The Manuscript Has “Keywords” Without a Dictionary
This is one of the most beautiful ideas in undeciphered-text research.
Meaning can leave distributional shadows.
If a book contains real topics, words associated with those topics may cluster even if the words themselves are unreadable. A term for “root” might prefer plant pages. A term for “month” might prefer zodiac pages. A personal name might cluster in one story. We do not know which Voynich tokens have those meanings, but we can search for the clustering behaviour.
The problem is that non-semantic mechanisms can cast similar shadows.
Scribe preferences cluster by pages. Encoding modes cluster. Source exemplars cluster. Locally generated word families cluster. Section-specific layout constraints cluster.
So a statistical keyword says:
This form carries information about region or context.
It does not yet say:
This form means “root”.
That gap is exactly where circular decipherments are born.
The Best Statistical Finding Is Often a Constraint, Not a Translation
This is worth saying slowly.
Suppose analysis shows that a token family occurs overwhelmingly in Currier B, prefers paragraph interiors, is absent from labels and clusters on Quire 13 pages.
That is valuable.
It constrains any proposed meaning. A candidate translation should make sense in those environments and explain the absences elsewhere.
But if we name the token from the page illustration first and then reinterpret its distribution to support our name, the direction of inference has reversed.
Good Voynich work tries to keep these stages apart:
- Measure where the form occurs.
- Describe what environments distinguish it.
- Propose mechanisms capable of generating that distribution.
- Only then compare semantic candidates.
- Require the semantic candidate to predict new occurrences or exclusions.
The difference looks procedural.
It is the difference between discovery and confirmation bias.
Could Voynichese Be a Natural Language?
Yes.
That answer should be allowed to stand without immediately becoming “therefore it is”.
A natural-language account has real attractions.
- Long-range vocabulary organisation is compatible with topical language.
- Token families can resemble morphology.
- Frequency distributions and network measures show language-like properties.
- Different textual registers can explain Currier and label/prose differences.
- Medieval scripts and abbreviation systems can make familiar languages look unfamiliar.
It also has burdens.
- The strong character restrictions require explanation.
- Low conditional entropy under common transcriptions is unusual.
- Some local repetition and near-neighbour behaviour can be stronger than expected in ordinary prose.
- No proposed underlying language has produced a broadly accepted, reproducible manuscript-scale reading.
- Visible spaces may not align with ordinary lexical boundaries.
A natural-language theory is therefore alive.
It still needs a representation mechanism.
Could Voynichese Be Ciphertext?
Yes.
Again, possibility is not identity.
Simple one-to-one substitution has long struggled with Voynich’s unusual distributions. But “cipher” is a much larger design space than one plaintext letter becoming one ciphertext symbol.
- Homophonic substitution can give one plaintext unit several ciphertext forms.
- Verbose substitution can expand one plaintext unit into multiple visible symbols.
- Nulls can insert non-semantic material.
- Nomenclators can encode frequent words or names separately.
- Stateful rules can change output depending on position.
- Abbreviation and ciphering can coexist.
Modern constructions such as the 2025 Naibbe cipher show that historically plausible hand-executable mechanisms can reproduce more Voynich-like surface structure than naïve substitution models do.
The burden remains formidable: a real cipher account must produce a consistent plaintext and explain the manuscript’s local and global structures rather than merely imitate statistics.
Could Voynichese Be Generated Text?
Yes.
Rule-based generation can reproduce several famous properties: repeated families, local similarity, constrained glyph ordering and long-tailed token distributions. Historical mechanisms such as combinatorial tables or grille-like methods have been proposed; modern algorithms can create remarkably Voynich-like surfaces.
But “generated” has two meanings that should not be confused.
Generated and meaningless
The visible strings are produced mainly to look text-like, without a recoverable underlying message.
Generated from meaning
An encoding, abbreviation, cipher or notation algorithm transforms meaningful source material into a highly constrained visible form.
The surface can look “generated” in both cases.
Therefore:
Generated-looking ≠ meaningless.
Could the Categories Be Wrong?
This may be the most mature answer available.
Natural language.
Cipher.
Abbreviation.
Artificial language.
Mnemonic notation.
Generated pseudo-text.
These categories are useful to us because they organise hypotheses.
A fifteenth-century maker was under no obligation to choose only one.
A text can be meaningful natural language written through heavy abbreviation. It can then be enciphered. Technical labels can use nomenclators while prose uses another scheme. Diagrams can use symbols differently from running text. A workshop can change convention across sections.
The true mechanism may therefore sit between our modern boxes.
This is not permission to invent arbitrarily complicated hybrid solutions.
Complexity must be earned by evidence.
But we should not reject a historically plausible mixed mechanism merely because it refuses our tidy taxonomy.
DON’T GO IN CIRCLES: The Statistics Edition
“Voynich follows Zipf-like frequency behaviour, so it must be language.”
No. Long-tailed frequency distributions are compatible with language but not unique to language.
“The text is non-random, so it must contain ordinary plaintext.”
No. Ciphers, structured generators, notation systems and abbreviation systems are non-random too.
“The entropy is too low for language, so it must be fake.”
No. Visible predictability can be introduced by the writing or encoding layer. The mechanism must be investigated rather than inferred from one statistic.
“This token is a statistical keyword, so it means the picture beside it.”
No. A keyword can classify context without revealing semantic value.
“These words differ by one glyph, so they are grammatical forms of one stem.”
Maybe. Local generation, orthographic variation, cipher expansion and abbreviation can also produce near-neighbour families.
“My generator reproduces Voynich-like words, so Voynich was generated this way.”
No. Reproducing selected properties demonstrates compatibility. Historical identity requires independent predictions and manuscript-scale fit.
“A cipher can now reproduce several properties, so Voynich is ciphertext.”
No. The 2025 Naibbe result keeps a cipher family viable; it does not identify the historical mechanism used in MS 408.
“A statistically language-like corpus tells us which language.”
No. Compatibility with natural-language structure and identification of the underlying language are different evidentiary tasks.
What Part V Leaves Standing
- Voynich tokens form dense families of near-related visible forms.
- Token interiors are strongly constrained; signs behave as if they occupy preferred structural zones.
- Local repetition and local vocabulary are genuine features of the corpus.
- Long-range token distributions contain organisation beyond simple shuffled randomness.
- Statistical studies have repeatedly found properties compatible with natural-language organisation.
- Those properties do not uniquely identify natural language, plaintext or a specific language.
- Character-level predictability is unusually strong under common transcriptions and must be explained by any serious mechanism.
- Transcription and segmentation choices affect entropy, length and combinatorial measurements.
- Historically plausible cipher constructions can reproduce multiple Voynich-like statistical properties from meaningful plaintext, so some older blanket arguments against ciphertext are too strong.
- Generated-text mechanisms can reproduce several surface properties without proving that Voynich is meaningless.
- No tested mechanism currently earns identity merely because it reproduces a subset of the statistics.
Structure is evidence. Structure is not translation. Statistical resemblance is evidence. Statistical resemblance is not mechanism identity.
Part V Reading Map
- Word Families
- Why Voynichese Is So Predictable
- Keywords Without Meanings
- The Q-Series
- Currier A and B
- What a Real Voynich Decipherment Must Survive
From Statistics Back to the Page
At this point it would be easy to let Voynich become a spreadsheet.
That would be another mistake.
Every token came from somewhere on a physical surface. Some sit beside roots. Some sit inside circles. Some label human figures. Some wrap around diagrams. Some begin lines. Some end paragraphs. Some stand beside stars.
The statistical structure and the visual structure belong to the same artefact.
The next frontier is therefore cross-modal.
Part VI — What the Pictures and Text Do Together asks the question that every plant-name guess, zodiac label and diagram caption has been trying to answer prematurely:
Can image and text constrain one another without either side being translated first?
Part VI — What the Pictures and Text Do Together
There is a bad habit Voynich encourages because the manuscript itself seems to cooperate with it.
We separate the pictures from the writing.
Botanists inspect plants. Cryptanalysts inspect strings. Art historians inspect iconography. Linguists inspect token distributions. Codicologists inspect quires. Each discipline can learn something real.
But the medieval maker did not receive those pages from five academic departments.
The writing and the pictures occupy the same parchment.
They avoid each other, surround each other, interrupt each other, label each other, cross folds together and sometimes compete for the same physical space. Whatever Voynich ultimately means, image and text were made to coexist as one information surface.
That gives us a frontier more powerful than asking pictures to translate words or words to caption pictures.
Can two unread systems constrain one another?
Yes.
They can.
And the reason is simple: relationships can be observed before meanings are known.
The Page Is the First Unit of Joint Evidence
Start with the page itself.
A plant is not pasted onto a completed paragraph. The prose bends around it. Lines stop where the image occupies space. A label may sit beside a root. On circular pages, writing curves around a diagram. On Quire 13, paragraphs occupy zones left open by tubes and figures. On the Rosettes foldout, blocks of text live inside a geometry that is larger than any normal page.
This means layout contains production information.
Even when local sequencing remains uncertain, a page can tell us which elements were designed to share one visual field. That relationship is stronger than two pages merely sitting next to each other in the present binding.
The hierarchy of evidence therefore looks something like this:
- same physical mark or continuous drawing;
- same local page architecture;
- same bifolium or continuous foldout relationship;
- same quire;
- present page adjacency;
- similar appearance elsewhere;
- external resemblance to another manuscript.
That order is not absolute, but it reminds us that internal physical relationships generally deserve examination before distant visual analogies.
Text Routing Is Evidence of Planning
Where writing stops is often as informative as where it begins.
When a line terminates at an image boundary, we learn that text production responded to the illustrated field. When prose occupies deliberately left blank zones, we learn that image and text were coordinated. When a label sits inside a diagram rather than in the surrounding prose, we learn that the scribe treated that string differently.
The sequencing question then becomes empirical.
Was the drawing made first and the prose fitted around it?
Were spaces reserved for pictures during writing?
Were image and text planned together from an exemplar?
Different folios may not have the same answer.
The important point is that page geometry can constrain production models without decipherment.
Labels Give Us the Tightest Image–Text Attachments
A long paragraph beside a plant can discuss almost anything related to the plant—or perhaps something not directly related at all.
A short string written directly beside one small star, root fragment, vessel, tube or human figure is more constrained.
It may name the object.
It may classify it.
It may assign a number, quality, status, date or role.
It may be an instruction directed at the object rather than its name.
But the space of plausible functions is narrower than for unrestricted prose.
This makes labels one of the most valuable bridges between visual and textual analysis.
Alignment can narrow semantic role without assigning semantic value.
A Label Beside a Plant Is Not Automatically the Plant’s Name
This rule deserves its own heading because entire decipherments have been built by ignoring it.
Imagine a modern diagram of a flower.
A string beside the drawing could say:
- rose;
- root;
- poisonous;
- harvest in spring;
- specimen 14;
- female plant;
- use two parts;
- from Padua;
- see page 27;
- do not boil.
Proximity tells us the text is probably associated with the object.
It does not identify the relation.
This is why “plant identification + adjacent token = dictionary entry” is one of the oldest circular paths in Voynich.
The Better Test: Does the Association Recur?
One image–token pairing can happen by chance or by flexible interpretation.
Repeated pairings are different.
Suppose a distinctive root motif appears on one large herbal page and again as a fragment on a later page. If the associated texts also share a rare token family, the visual and textual evidence begin to converge.
That still does not tell us whether the shared string means the species, the root, a property, a preparation or something else.
But it does establish a stronger internal relationship.
Now imagine the same pattern across ten independent examples.
That becomes a serious structural clue.
The general principle is:
A cross-modal relationship becomes valuable when it generalises beyond the example that suggested it.
Internal Recurrence Is Stronger Than External Resemblance
This is where Voynich can become its own reference dictionary before it becomes a readable dictionary.
If the same unusual visual motif recurs within the manuscript, we can ask whether its textual neighbourhood also recurs. We do not need to know what the motif is called in Latin, Italian or any modern botanical catalogue.
The manuscript can tell us:
- these two images are unusually similar;
- these two labels share a component;
- these two page contexts belong to the same Currier regime;
- this token family clusters near this image family;
- this association does not occur in a contrasting section.
Only after establishing that internal relation do we ask what external semantic category best explains it.
This reverses the usual popular method, which begins with an external object name and then looks inward for confirmation.
The Zodiac Pages Offer a Rare Controlled Context
Zodiac pages are especially useful because the central sign can often be identified independently of Voynichese.
That does not decode the surrounding labels, but it gives us an externally constrained visual environment.
Suppose labels around Pisces differ systematically from labels around Aries, Taurus and Gemini. That could reflect dates, stars, degrees, people or another zodiac-related subdivision. Suppose instead the label vocabulary is largely stable across signs. That would favour a different class of explanations.
The point is not to guess which answer feels right.
The point is that a known central icon gives us a place to measure conditional variation.
The zodiac therefore offers something close to a natural experiment:
known visual class, unknown textual function.
Quire 13 Gives Us a Different Controlled Context
Quire 13 repeatedly combines figures, containers, connections and labels.
Here the visual semantics are less externally anchored than the zodiac, but the repeated document grammar is stronger.
We can ask whether labels attached to human figures behave differently from labels attached to tubes or vessels. We can compare their token lengths. We can ask whether similar diagram positions receive related strings. We can test whether the same label families appear in running prose. We can ask whether a label at a junction behaves differently from one inside a pool.
Even a negative result matters.
If labels attached to apparently similar visual units show no systematic textual relationship, a simple naming model weakens.
If they do show systematic relationships, the next question becomes whether that relationship is semantic, grammatical, positional or generated by another mechanism.
The Rosettes Foldout Forces Us to Preserve Geometry
A conventional transcription turns a spatial surface into a sequence of lines.
That is necessary for many forms of analysis.
It is also lossy.
On the Rosettes foldout, text belongs to zones, bands, circles and connections. A string’s relationship to the centre, edge, pathway or neighbouring rosette may be as important as its place in a linear transcription file.
A future model that treats every token only as position 14,287 in a character stream may therefore discard precisely the information the medieval page was built to preserve.
This is why serious computational representation should eventually include:
- folio coordinates;
- orientation;
- line and paragraph identity;
- label attachment;
- diagram region;
- distance from illustrated objects;
- connectivity across foldout regions;
- bifolium relationships;
- uncertainty about reading order.
Voynich is not merely a string problem.
It is a spatial document problem.
The Same Shape Can Do Different Jobs
Stars are a perfect example.
Star-like forms occur in several contexts.
- as apparent celestial objects in circular diagrams;
- held by human figures;
- as repeated marginal markers in Quire 20;
- as decorative or structural motifs in other visual fields.
It would be tempting to insist that “star always means star”.
But writing systems and diagrams reuse shapes functionally.
A modern asterisk can mark a footnote, multiplication, wildcard, correction, emphasis or a rating. Its form remains recognisable while its role depends on context.
Voynich may do something similar.
Same shape ≠ same function.
That rule protects us from building a universal symbolic dictionary from cross-page resemblance alone.
Text Can Constrain a Picture Without Naming It
Suppose two visually similar objects receive sharply different textual treatment.
One always has a short label. The other always sits inside prose. One appears almost exclusively in Currier A. The other belongs to Currier B. One is associated with a narrow token family. The other has no such relationship.
That difference tells us the pictures may not belong to the same functional class despite visual resemblance.
The reverse is also true.
If visually different objects repeatedly receive text with the same distinctive structural pattern, they may share a functional category invisible to us.
This is where cross-modal analysis becomes more than decoration. Text can partition the images. Images can partition the text.
The Whole-Plant → Fragment Relationship Remains Promising but Incomplete
One of the most attractive internal hypotheses in the manuscript is that some small pharmaceutical-looking plant fragments correspond deliberately to components of large herbal plants shown earlier.
There are visually suggestive cases.
What would make the relationship strong?
- distinctive morphology, not merely generic leaf or root shapes;
- multiple independent matches;
- consistent label relationships;
- shared rare token components beyond chance expectation;
- a stable rule for which part of the whole plant is extracted;
- similar treatment of matched pairs across different hands or Currier regimes;
- predictive success on cases not used to formulate the matching rule.
The hypothesis remains valuable precisely because it can be made harder to satisfy.
Easy theories explain everything after the fact.
Good theories become less flexible as evidence accumulates.
The Grand Pipeline Failed Because It Was Too Good a Story
We have already frozen this result in Part II, but it belongs here for a deeper reason.
The proposed chain was irresistibly coherent:
identify plant → select part → prepare/store → apply to body → time by stars → record recipe
It connects nearly every visual regime.
That is exactly why it became dangerous.
A beautiful explanatory chain can absorb ambiguity at every step. A plant-like page becomes identification because the theory needs identification. A vessel becomes preparation because the theory needs preparation. Bodies become application because the theory needs application. Zodiac becomes timing because the theory needs timing. Starred entries become recipes because the theory needs recipes.
Once the story is installed, every page appears to confirm it.
Our broader work found that several individual relationships remain historically plausible and structurally interesting, but the universal sequence did not survive strongly enough across the whole manuscript to earn master-architecture status.
Why not?
- current page order cannot be treated as guaranteed process order;
- not every image family supplies the predicted cross-text relationship;
- visual section boundaries and textual regimes do not align neatly enough;
- label behaviour is richer than a simple noun inventory;
- some local structures appear autonomous rather than stages in one global process;
- too many semantic steps had to be inferred from visual resemblance before independent textual support existed.
So we keep the pieces.
We stop pretending the pieces already form the whole machine.
Current Adjacency Is Not Process Order
This one rule destroys many elegant Voynich narratives.
A page of plants followed by a page of containers does not prove “plants are put into containers”. A zodiac page followed by Quire 13 does not prove “astrology controls bathing”. A fragment page before starred text does not prove “ingredients become recipes”.
Present sequence may preserve some intended order.
It may also reflect lost bifolia, rebinding or gathering-level changes.
A process-order claim therefore needs more than adjacency.
- repeated transitions;
- cross-reference-like text;
- physical continuity;
- predictable object transformations;
- systematic label relationships;
- or another independent signal that sequence is functional.
Current order ≠ original order ≠ process order.
What Would a Genuine Cross-Modal Breakthrough Look Like?
It would probably look less spectacular at first than a viral translation.
Imagine discovering that one rare token component occurs in labels attached to a particular visual feature across ten unrelated folios, and that it is absent when that feature is absent.
That would not tell us the English meaning.
But it might identify a stable functional relationship.
Now suppose the relationship predicts an eleventh case on a folio nobody used to construct the hypothesis.
Now we have something much stronger.
Only after that would we compare external semantics.
This is slower than saying “this label means star”.
It is also how the field might eventually stop restarting itself.
The Missing Negative Cases Matter
Suppose a proposed image–text rule succeeds on five examples.
That sounds impressive.
What if the same visual feature occurs twenty more times and the textual rule fails on eighteen?
The five successes no longer mean what they seemed to mean.
This is why negative cases must be preserved.
Voynich is especially vulnerable to selective success because the corpus is large enough to provide many coincidences. A determined reader can nearly always find several pages that support an idea.
The correct denominator is not “how many examples fit?”
It is “how many eligible examples were there, and what happened to all of them?”
We do not expose the private machinery used in our own investigations, but the public lesson is simple:
A pattern that survives only because failures are forgotten is not a surviving pattern.
What Part VI Leaves Standing
- Text and image are materially integrated on many Voynich pages.
- Text routing and reserved space provide evidence about production planning.
- Labels create unusually tight image–text relationships while remaining semantically ambiguous.
- A label’s spatial attachment can constrain role without proving that it is a noun or name.
- Repeated internal image–text associations are more valuable than isolated external resemblance.
- Zodiac pages provide known visual classes against which unknown labels can be compared.
- Quire 13 provides repeated structural classes—figures, containers, tubes, labels—that can be compared without pre-decoding them.
- The Rosettes foldout requires preservation of spatial geometry, not merely linear transcription.
- The same visual form, such as a star, can serve different functions in different contexts.
- Whole-plant ↔ fragment relationships remain a promising internal research direction but are not yet a solved semantic bridge.
- The universal plant → part → preparation → body → astrology → recipe pipeline did not generalise strongly enough to become established architecture.
- Current adjacency cannot substitute for demonstrated process sequence.
- Negative cases are part of the evidence and must travel with positive cases.
DON’T GO IN CIRCLES: The Image–Text Edition
“The label is beside the object, so it names the object.”
It is associated with the object. Naming is only one possible relation.
“The same star shape appears twice, so it has one universal meaning.”
Form can recur across functions. Context decides whether a cross-regime equivalence is justified.
“This plant fragment resembles that whole plant, so the relationship is solved.”
Visual similarity is the beginning. Repeated morphology, labels, vocabulary and independent prediction are what make the relationship strong.
“These consecutive sections describe consecutive operations.”
Not without independent evidence. Binding history, missing leaves and non-aligned textual regimes make present adjacency an unsafe substitute for workflow.
“I found five matches.”
How many eligible non-matches were there? Voynich rewards theories that remember their denominator.
The Next Step Is Historical Comparison
Once internal relationships are mapped, we can finally ask the external question more safely.
What kinds of books existed around Voynich’s material horizon?
How did other manuscripts combine plants and prose? How did they label ingredients? How did medieval medicine use zodiac diagrams? How were baths represented? How were recipes segmented? How did scribes compress knowledge when words, pictures and page geometry had to share expensive parchment?
This is where Carrara, Padua, Masson 116, Sloane 4016, Egerton 747 and other comparators become valuable.
But only if we remember the rule before entering the library:
Comparator is not source. Resemblance is not descent.
Part VII — The Manuscript World Around Voynich begins there.
Part VII — The Manuscript World Around Voynich
Voynich does not have to resemble another manuscript in order to belong to history.
That sentence sounds obvious. It is not how the subject usually behaves.
The moment somebody finds a medieval herbal with a strangely shaped root, a bath manuscript with naked bodies, an astronomical diagram with concentric circles or an Italian pharmacy jar with an ornate neck, the gravitational pull begins.
This looks like Voynich.
Therefore this may be where Voynich came from.
Therefore this may tell us what Voynich means.
Three different claims have just been compressed into one emotional reaction.
Part VII exists to pull them apart again.
Comparator is not source. Resemblance is not descent.
A comparator can still be enormously valuable. In fact, the right comparator often teaches us more than an attempted “match”. It shows what was technically possible, what kinds of information medieval makers commonly placed together, how pictures and words shared a page, how copies changed over generations and which visual habits belonged to a wider knowledge culture rather than to Voynich alone.
The aim is not to find Voynich’s twin.
The aim is to reconstruct the neighbourhood of possible books around it.
What a Good Comparator Can Actually Prove
Before opening another manuscript, decide what job the comparison is allowed to do.
Historical feasibility
If another manuscript from a relevant period combines medicinal plants, astronomical timing and recipes, then such a combination was historically possible. That matters if somebody claims those subjects could never have coexisted in one learned book.
Information architecture
A comparator can show how medieval scribes arranged full-page images, captions, ingredient names, diagrams, indexes, lists, cross-references and prose. It teaches us what page structures existed.
Iconographic convention
If roots, vessels, zodiac signs or bathing bodies have conventional ways of being drawn, comparison helps us distinguish culturally familiar motifs from features that may be unusually Voynich.
Copying behaviour
Manuscripts that survive in known relationships show how images drift. A copy may simplify one leaf, exaggerate another root, change colour, omit text, add text, reorder material or preserve a recognisable prototype despite substantial visual change.
Knowledge ecology
Collections reveal which subjects medieval readers and compilers allowed to share one intellectual environment. A medical codex containing an herbal, an antidotary, substitutions, weights, measures and a lunar calendar tells us something about how knowledge was bundled.
Those are powerful uses.
Now the limits.
What a Comparator Cannot Prove by Looking Similar
- It cannot prove that the Voynich maker saw that manuscript.
- It cannot prove direct copying.
- It cannot prove a shared workshop.
- It cannot prove a city of origin.
- It cannot prove the language of Voynichese.
- It cannot prove that similar-looking pictures served the same function.
- It cannot prove that a visual resemblance was historically transmitted rather than independently conventional.
- It cannot convert a plausible region into provenance without a custody, textual, palaeographic or material bridge.
The difference between “this makes sense in the same world” and “this came from this source” is the difference between context and genealogy.
Voynich needs much more context.
It also needs much less imaginary genealogy.
The Carrara Herbal: Why It Matters So Much
The Carrara Herbal is one of the most useful comparators in the entire eduKate Voynich investigation because it sits close to the right historical horizon while being well enough understood to stop us inventing its function.
Today it is British Library Egerton MS 2020. The British Library describes it as an illustrated herbal containing a fourteenth-century Italian translation by Jacopo Filippo of a work associated with Serapion the Younger. The manuscript is usually dated around 1390–1404. Its decoration includes Carraresi heraldry connected with Padua and numerous coloured plant miniatures. British Library catalogue record.
That date matters because it places the Carrara Herbal immediately before the radiocarbon window associated with Voynich parchment.
The intellectual relationship is equally interesting. The Carrara Herbal is not a modern field guide in which one plant image equals one modern species caption and nothing more. It belongs to a medicinal textual tradition transmitted across languages and centuries. Arabic pharmacological knowledge, Latin mediation, vernacular translation, manuscript illustration and local patronage meet in one object.
That alone teaches us something crucial.
A medieval plant manuscript can be simultaneously botanical-looking, medical, translated, copied, adapted and visually reorganised.
When a Voynich plant looks strange to a modern botanist, we should therefore not assume the maker lacked botanical knowledge. Nor should we assume the image was meant primarily to satisfy modern species-recognition criteria. The plant can be serving a medicinal, mnemonic, textual or copying function inherited from another representational culture.
Carrara Shows What Voynich Does Not Give Us
This is just as valuable.
In Carrara we have readable language. We have known textual tradition. We can compare the images with the plant names and medicinal discussion. We can follow scholarly identification work. We have a cultural environment that is historically legible.
Voynich removes those anchors.
That means Carrara should not be used to fill the gaps by imagination. It should be used to show which kinds of relationships we would want to recover.
For example:
- image ↔ name;
- name ↔ synonym;
- plant ↔ medicinal property;
- plant ↔ preparation;
- vernacular term ↔ inherited scholarly tradition;
- image ↔ exemplar lineage.
Those relationships are demonstrable in a deciphered manuscript.
They remain hypotheses in Voynich until the internal evidence supports them.
That is why our Carrara work became more useful as a control than as an answer.
Padua Is a Knowledge Environment, Not a Proven Birth Certificate
Padua attracts Voynich theories for understandable reasons.
Late medieval northern Italy supported universities, physicians, apothecaries, artists, courtly patronage, vernacular translation, astrology, manuscript production and botanical traditions. Padua in particular became a powerful centre of medical learning. The Carrara Herbal sits within that intellectual landscape.
This makes Padua an excellent possibility environment.
It does not make Padua the demonstrated birthplace of MS 408.
The distinction becomes especially important when several weak similarities accumulate. An Italian-looking plant. A bath tradition. A zodiac diagram. A jar form. A manuscript from nearby decades. A plausible university milieu.
Ten plausible contextual correspondences can still add up to context rather than provenance.
To move from “northern Italy could host such a manuscript” to “Voynich was made in Padua”, we would need a much sharper bridge: a distinctive scribal tradition, document trail, identifiable exemplar, local material feature, securely diagnostic iconography, linguistic evidence or another independent regional marker.
Historical fit ranks candidates. It does not sign the colophon.
Masson 116 and Sloane 4016: Copying Is Transformation
Another important lesson comes from a known manuscript relationship.
The British Library catalogues Sloane MS 4016 as an Italian herbal of about 1440, part of a North Italian group, and describes it as a copy of Paris, École des Beaux-Arts MS Masson 116. Its pages contain large coloured plant miniatures, often accompanied by people or animals and captions. British Library catalogue record for Sloane MS 4016.
Why should Voynich researchers care?
Because here we have something Voynich speculation often lacks: an actual copying relationship discussed in manuscript scholarship.
Once we know a copy relationship exists, we can study what changed.
A plant does not need to remain photographically identical across copies. The human hand simplifies, stylises, regularises and sometimes misunderstands. A leaf may become more symmetrical. A root may become more dramatic. A missing region may be reconstructed. A decorative tradition can migrate into a practical diagram. Captions can survive while explanatory text changes. Text can be abbreviated while the picture remains expansive.
This is a powerful corrective to two opposite mistakes.
Mistake one: “The Voynich plant does not look naturalistic, so it cannot come from a real herbal tradition.”
Copying history can distort morphology dramatically while preserving enough structure for a manuscript tradition to remain recognisable.
Mistake two: “This Voynich plant resembles one medieval herbal image, so that manuscript must be the source.”
A genuine source relationship requires a family of shared errors, distinctive correspondences, sequencing, textual echoes or other genealogical signals. One visual resemblance is not enough.
Known copying traditions teach us how demanding a real descent argument should be.
The Roccabonella Herbal: Even a Related Image Tradition Can Change
The Roccabonella Herbal, usually identified with Venice, Biblioteca Nazionale Marciana, MS Lat. VI, 59 (= 2548), provides another useful control. Scholarship on medieval plant iconography dates it to around the middle of the fifteenth century and describes its illustrations as drawing from the Carrara Herbal tradition while introducing changes. A botanical-historical study of medieval Cucumis imagery, for example, explicitly notes that the Roccabonella images reproduce Carrara material with modifications. See the comparative study.
This gives us a miniature lesson in manuscript evolution.
An image tradition can remain related while individual images drift.
Therefore exact visual identity is not required for descent.
But neither does visual drift make every vaguely similar plant part of the same lineage.
The useful question is not “Does this look a bit like that?”
It is “Can we recover a transformation path across multiple independent details?”
Egerton MS 747: A Medieval Medical Book Did Not Need One Subject
British Library Egerton MS 747 is useful for a completely different reason.
Its catalogue describes a collection of medical texts containing the Tractatus de herbis, additional plant images, an added forecasting text with a lunar calendar, the Antidotarium Nicolai, material on doses, a list of substitutions for unavailable ingredients, texts on weights and measures, and synonyms for plant names and medical ingredients. The British Library identifies it as the oldest surviving copy of the Tractatus de herbis. British Library catalogue record.
Read that list again slowly.
Plants.
A lunar calendar.
Medicinal compounds.
Doses.
Ingredient substitutions.
Weights and measures.
Synonyms.
A modern classifier might be tempted to place each subject in a different folder. The medieval codex did not have to.
This is one reason the Voynich combination of plant imagery, celestial-looking diagrams, vessels and short entries should not be dismissed as automatically incoherent.
Medieval practical knowledge could be bundled.
But Egerton 747 does not prove that Voynich is a medical miscellany. It proves that such cross-domain bundling was historically real.
A comparator can defeat an argument from impossibility without proving the Voynich alternative.
This Changes How We Think About “Sections”
If a medieval medical codex can move among herbals, lunar timing, antidotaries, dosage, substitutions and synonyms, then the existence of several visual regimes inside Voynich no longer demands five unrelated books accidentally bound together.
They could be unrelated.
They could be related.
The comparator simply tells us that thematic breadth is not historically absurd.
Notice again how mature comparison narrows certainty rather than inflating it.
Medieval Medicine and the Sky Could Share a Table
The presence of zodiacal imagery beside plant and bodily material often feels strange because twenty-first-century knowledge has separated astronomy, astrology and clinical medicine into different epistemic categories.
That separation cannot simply be projected backward.
Medieval and Renaissance medicine could use celestial timing in relation to bloodletting, regimen, diagnosis and treatment. Lunar calendars could sit inside medical collections. Zodiacal-man imagery could relate regions of the body to signs. Physicians could practise learned medicine while accepting celestial influence as part of natural philosophy.
Therefore a Voynich manuscript that combines celestial and medical-looking material would be historically conceivable.
The important phrase remains historically conceivable.
We should not quietly upgrade it to “therefore the zodiac pages schedule the recipes”.
Balneology Was a Real Technical Literature
Quire 13 becomes much more interesting when compared with the history of therapeutic bathing—not because the comparison solves the images, but because it removes the idea that bodies in medicinal water would be culturally bizarre.
One famous tradition is Pietro da Eboli’s De balneis Puteolanis, a poem on the therapeutic baths around Pozzuoli and Baia. Surviving illuminated manuscripts pair descriptions of individual baths with scenes of bathing and treatment. The Biblioteca Angelica describes its manuscript 1474 as a luxury medieval witness to a work composed in the early thirteenth century celebrating the curative properties of Campanian thermal waters. Biblioteca Angelica overview.
Later medical writers treated bathing even more systematically. Ugolino da Montecatini’s Tractatus de balneis, associated with the early fifteenth century, discussed therapeutic waters, their properties and modes of treatment. Treccani records manuscript witnesses in Florence’s Biblioteca Medicea Laurenziana and Pavia’s university library and notes the work’s place in the development of medical hydrology. Treccani biography of Ugolino Caccini da Montecatini.
By the fifteenth century, bathing was therefore not merely an everyday action waiting to be drawn. It could be an organised subject of medical writing.
This keeps a balneological interpretation of some Voynich imagery plausible.
It also shows us what a real bath tradition tends to contain: named places, described waters, conditions, procedures and textual explanation.
Voynich Quire 13 still has to demonstrate its relationship to those functions rather than inheriting them from resemblance.
The Baths of Pozzuoli Teach Another Lesson: Copies Can Preserve an Old Visual System
BnF records for De balneis Puteolanis show a manuscript tradition extending across centuries. A fourteenth-century witness can preserve an iconographic structure derived from earlier models; the BnF notes that many surviving copies belong to a southern Italian tradition spanning the thirteenth to fifteenth centuries. Biblissima/BnF manuscript record.
This is another warning against dating imagery purely by how “old-fashioned” it looks.
A fifteenth-century copy can deliberately preserve forms originating much earlier.
Voynich imagery therefore needs to be dated through the object and comparative art history together, not by assuming every visual convention was invented at the moment the parchment was written.
Pharmacy Jars: A Useful Material Control, Not a Caption
The ornate Voynich vessels inevitably invite comparison with apothecary containers.
Real late medieval and Renaissance pharmacies did use repeated storage vessels in standardised families. The Metropolitan Museum of Art describes Italian maiolica storage jars as common pharmacy or shop equipment, often designed to stand in sets on shelves, with forms adapted for gripping and sealing and sometimes carrying inscriptions identifying contents. Surviving Italian examples from the late fifteenth century show that differentiated pharmacy vessels were materially real, not an invention of modern Voynich interpretation. The Met: Italian pharmacy jar, c. 1475–1500.
This gives the Voynich vessel comparison a legitimate material basis.
It still does not make a drawn vessel an albarello by decree.
Several questions remain:
- Do the Voynich forms reproduce construction features of real storage vessels or only the general idea of a decorated container?
- Do different vessel shapes correlate with different textual or plant-fragment classes?
- Are the differences functional, decorative or symbolic?
- Do labels behave like contents labels?
- Does the visual vocabulary match a particular regional ceramic tradition strongly enough to discriminate geography?
The first answer may be “possibly”.
The later answers are where proof would live.
Why “Looks Italian” Is So Dangerous
Once Carrara, Roccabonella, Padua, Italian baths and Italian pharmacy vessels enter the discussion, a reader may feel the case tightening toward Italy.
That feeling has to be audited.
Are we collecting independent regional indicators?
Or are we repeatedly selecting Italian comparators because we already suspect Italy?
The second process can create an illusion of convergence.
If we search only Italian libraries, every Voynich feature will eventually acquire an Italian cousin. Search German, French, Bohemian, Iberian, Byzantine, Arabic and Hebrew manuscript traditions with equal effort and the exclusivity of the original match may disappear.
A regional hypothesis becomes strong not when examples can be found there, but when the same diagnostic combination is difficult to find elsewhere.
Similarity becomes geographical evidence only when alternatives are systematically less similar for reasons that matter.
A Comparator Should Have a Job Before We Look at It
This is one of the simplest ways to stop comparison from becoming treasure hunting.
Before opening the comparator, write down the question.
- Can a medical manuscript combine lunar material with pharmacy?
- How do known herbal copies transform roots?
- How are therapeutic baths represented?
- What does an actual labelled apothecary vessel look like?
- How do circular diagrams organise captions?
- How much text can disappear while image lineage remains recognisable?
- Which page relationships indicate copying rather than shared convention?
Then inspect the comparator for that job.
Do not let every interesting feature become a new proof of your favourite origin theory.
A manuscript can teach us three unrelated things without becoming Voynich’s parent.
The Most Useful Comparator May Be the One That Looks Less Like Voynich
This sounds backwards.
Suppose two books look very similar. We learn that visual resemblance exists.
Suppose a third book looks different but performs the same information task.
Now we learn how many visual solutions that task can have.
That is often more valuable.
A medical recipe does not need one universal layout. A plant label does not need one position. A zodiacal calendar does not need one ring design. Bathing can be depicted architecturally, narratively or schematically. Storage can be represented by actual vessel shapes or abstract containers.
Functional variation helps prevent us from mistaking one aesthetic convention for the only possible implementation of a medieval task.
Masson 116 Teaches a Powerful Absence Lesson
One reason Masson 116 became important in our wider work is that manuscript survival is uneven.
A manuscript can preserve images whose textual context is incomplete, altered or difficult to reconstruct. A later copy can preserve visual information that changes our understanding of an earlier witness. Missing captions do not automatically mean the images were originally meaningless. Surviving captions do not guarantee that all original explanatory text survives.
This matters for Voynich because we repeatedly face the temptation to infer original function from surviving layout as if nothing had been lost.
But Part I already established missing leaves.
Known manuscript traditions show that loss and copying can selectively preserve different information layers.
Therefore absence must be handled symmetrically:
Missing evidence may explain uncertainty. It cannot be used as positive evidence for whatever we wish the missing page contained.
Knowledge Travelled Across Languages
The Carrara Herbal is also a reminder that medieval knowledge was already multilingual.
A work associated with an Arabic pharmacological tradition could move through Latin mediation into an Italian vernacular manuscript with new illustrations, local patronage and its own copying history.
That makes simplistic linguistic provenance dangerous.
A manuscript produced in northern Italy does not have to preserve knowledge originating in northern Italy.
A vernacular text can contain Arabic-derived material.
A Latin scribe can copy a Greek authority.
A visual tradition can travel farther than its language.
An encrypted text can conceal a language different from the place of copying.
So even if Voynich were eventually located geographically, that would not automatically identify its intellectual ancestry or underlying language.
Origin, Discovery, Use, Custody and Intellectual Ancestry Are Different
This distinction is essential far beyond Voynich.
- Origin: where the physical codex was produced.
- Intellectual ancestry: where the knowledge traditions behind it came from.
- Use: where and by whom the codex was actually consulted.
- Custody: who owned or stored it at different moments.
- Discovery: where a modern researcher encountered it.
- Current location: where the object is held today.
Wilfrid Voynich acquired the manuscript in Italy in 1912.
That does not mean it was made in Italy.
Marci sent it from Prague to Kircher in Rome in the seventeenth century.
That does not mean it was made in Prague or Rome.
A possible northern Italian iconographic affinity would not mean its underlying knowledge was locally invented.
Each claim needs its own evidence.
A Known Copying Chain Is Gold
When we know that one manuscript derives from another, the pair becomes a laboratory.
We can ask what copying preserves and what it destroys.
Does root topology survive?
Does leaf count survive?
Does colour survive?
Do captions move?
Are species identities preserved more reliably than proportions?
Does a copyist “improve” an unfamiliar plant into a familiar one?
Does one tradition split a composite figure into parts?
Those empirical transformation rules can then be applied cautiously to Voynich.
That is far stronger than saying “medieval artists were inaccurate”.
We want to know how they were inaccurate, under which copying conditions, and which features tended to remain stable.
A Good Comparator Should Make the Theory Harder
This is our preferred test.
If comparison only gives a theory more ways to be right, the comparison is being used badly.
Suppose somebody proposes that Voynich plant images descend from a particular northern Italian herbal family.
The comparator should now create obligations.
- Which diagnostic image features should recur?
- Which ordering relationships should survive?
- Which plant pairs should be recognisable?
- Which characteristic errors should be shared?
- Which features should not occur if another tradition is the source?
- What date range is compatible with transmission?
- What linguistic or geographical bridge is needed?
If the answer to every mismatch is “the Voynich maker changed that”, the comparator has stopped testing the theory.
It has become decoration.
What We Learnt From Carrara Without Making It Voynich
- A richly illustrated medicinal herbal existed immediately before the Voynich parchment horizon.
- Plant image, vernacular language and inherited pharmacological tradition could coexist in one codex.
- Plant images can be information-bearing without behaving like modern field-guide plates.
- A manuscript can mediate knowledge across Arabic, Latin and vernacular traditions.
- Known manuscript lineages show that images can change while relationships survive.
- Northern Italian manuscript culture is a historically plausible comparator environment.
- None of those facts proves Voynich was made in Padua, copied from Carrara or written in Italian.
That final bullet is not a disappointment.
It is the reason the earlier bullets remain trustworthy.
What We Learnt From Medical Miscellanies Without Calling Voynich Medicine
- Plants, drugs, substitutions, measures and lunar material could occupy the same medical knowledge environment.
- Medieval manuscript architecture did not require one modern academic subject per codex.
- Practical retrieval structures—lists, captions, entries and indexes—were part of learned manuscript design.
- Astrological or calendrical material beside medicine is historically plausible.
- None of this independently establishes Voynich’s function.
What We Learnt From Bath Manuscripts Without Calling Quire 13 a Spa Guide
- Therapeutic bathing was a serious medieval medical subject.
- Bath manuscripts could combine text with repeated human figures in water.
- Knowledge of individual waters, places and therapies could be organised into separate entries.
- Bath imagery has a real historical comparator family.
- Voynich Quire 13 remains more structurally networked and enigmatic than a simple identification with any known bath text can explain.
What We Learnt From Pharmacy Without Making Every Vessel an Albarello
- Differentiated storage vessels were normal features of late medieval and Renaissance pharmacy.
- Sets of vessels could be visually distinguished and labelled.
- Container form can encode practical information.
- Voynich’s vessel-like drawings therefore have a legitimate pharmaceutical comparison.
- The drawings still require internal textual and morphological evidence before a specific vessel function is assigned.
DON’T GO IN CIRCLES: The Comparator Edition
“This manuscript is from the right period and looks similar, so it is the source.”
No. Chronological compatibility is a prerequisite for descent, not proof of descent.
“Carrara is Paduan, so Voynich is Paduan.”
No. Carrara demonstrates a relevant manuscript culture. Voynich provenance requires independent localisation.
“The Voynich plants resemble North Italian herbals, so the script must be Italian.”
No. Image tradition, production location and underlying language are separable historical variables.
“A bath manuscript proves Quire 13 is balneology.”
No. It proves therapeutic bathing belongs to the historical possibility space.
“An apothecary jar resembles a Voynich vessel.”
Good. Now ask which structural details, label relations and vessel classes generalise beyond the resemblance.
“The same subject mix appears in another medical manuscript, so Voynich is medical.”
No. The comparator defeats the claim that the mix is historically impossible. Function remains to be demonstrated.
“I found a perfect visual match.”
Then search aggressively for non-matches and alternative traditions. A match becomes diagnostic only when its distinguishing features are rare enough to exclude competitors.
What Part VII Leaves Standing
- Voynich belongs to a medieval manuscript world in which plants, drugs, calendars, celestial influence, baths, recipes and practical medicine could coexist in overlapping knowledge systems.
- The Carrara Herbal gives a particularly strong near-contemporary control for illustrated medicinal herbal culture in northern Italy.
- Sloane 4016 and Masson 116 provide a known copying relationship useful for studying visual transmission rather than guessing it.
- The Roccabonella tradition shows related imagery can change across manuscript generations.
- Egerton 747 demonstrates that a medical codex could bundle herbal, lunar, antidotary, substitution, measurement and synonym material.
- De balneis Puteolanis and Ugolino da Montecatini demonstrate that therapeutic waters generated a genuine medieval textual and visual tradition.
- Surviving pharmacy vessels show that repeated, differentiated storage containers belonged to real medical material culture.
- These comparators expand and discipline the historical possibility space.
- None independently establishes Voynich’s source, city, language, author, function or direct manuscript lineage.
Use other manuscripts to make Voynich theories harder, not easier.
Part VII Reading Map
- The Carrara Herbal — the broader eduKateSG comparator work.
- Medieval Medical Manuscripts of Padua: From Book to Patient — knowledge environment, not claimed Voynich provenance.
- The Herbal Pages
- The Human Figures, Pools and Tubes
- The Vessels and Plant Fragments
- The Zodiac Pages
Now We Can Finally Talk About Failure Properly
For seven parts we have accumulated constraints.
Physical constraints.
Visual constraints.
Textual constraints.
Statistical constraints.
Historical constraints.
Now we can perform the most useful operation in the entire master.
We can stop carrying ideas that have already failed.
Not because failed ideas were foolish.
Because a failed hypothesis is valuable once its failure is remembered.
Part VIII — What We Tried, What Broke, and What We Should Stop Repeating is the anti-loop centre of this entire article.
Part VIII — What We Tried, What Broke, and What We Should Stop Repeating
This is the part of the article that matters most if you want Voynich research to move forward.
Not because it contains the solution.
Because it contains something the subject has been unusually bad at preserving.
Defeat.
A good hypothesis that fails is not wasted work. It is a wall we no longer need to walk into at full speed. It tells the next researcher which bridge did not carry weight, which resemblance stopped generalising, which attractive interpretation depended on an assumption that the manuscript refused to keep.
Voynich has a cultural problem as much as a cryptographic one: failed ideas do not disappear. They become new again when somebody encounters the manuscript without encountering the history of the failure.
A plant is identified again.
A city is found in the Rosettes again.
A language is read from three labels again.
A zodiac theory becomes a medical timing system again.
A statistical curve is declared impossible for anything except language—or impossible for language—again.
An algorithm creates Voynich-looking strings and is promoted from demonstration to historical mechanism again.
Then years later another reader begins at the same starting line.
A field that remembers only its exciting hypotheses and forgets why they failed cannot accumulate knowledge efficiently.
This section therefore does something unfashionable.
It keeps the failures.
How to Read a Failed Voynich Theory
We are not interested in humiliating people who proposed ideas. That would teach nothing.
Some failed theories were intelligent. Some generated useful datasets. Some noticed real structures before anyone else. Some failed only because later evidence became available. Some are not completely dead; they simply do not deserve the confidence once attached to them.
For each recurring idea, ask four things.
- Why was it tempting? What genuine feature of the manuscript made a reasonable person consider it?
- What survived? Which observation remains useful even after the larger theory weakens?
- What broke? Which prediction, comparison, chronology, distribution or wider-manuscript requirement did not survive?
- What would reopen it? What genuinely new evidence would justify returning to the idea rather than merely repeating it?
This is how negative knowledge becomes cumulative.
We will discuss the conclusions of our own investigations where they are useful to readers, but not the private testing machinery that produced them. You do not need our internal protocols to benefit from a public result such as “this relationship did not generalise across the eligible pages”.
The result is the public asset.
1. Roger Bacon Wrote the Voynich Manuscript
Why it was tempting
The Roger Bacon story is old enough to feel like part of the manuscript itself. Johannes Marcus Marci’s seventeenth-century letter to Athanasius Kircher reports that the book had been bought by Emperor Rudolf II for 600 ducats and that the emperor believed it to be a work of Roger Bacon. Wilfrid Voynich later promoted the Bacon association enthusiastically.
Bacon was an irresistible candidate: medieval, learned, interested in natural philosophy, optics, languages and secrecy, and surrounded in later imagination by an aura of forbidden knowledge.
What survived
Marci’s letter is genuine historical evidence that a Bacon attribution circulated in the manuscript’s seventeenth-century history. That tells us something about how the codex was interpreted and valued.
It does not tell us Bacon wrote it.
What broke
Roger Bacon died around 1292. Radiocarbon testing places sampled Voynich parchment in the early fifteenth century, commonly summarised as 1404–1438. The surviving codex therefore cannot straightforwardly be a physical manuscript penned by Bacon more than a century earlier.
Yale’s own modern account notes that the carbon dating disproved the old direct Bacon-authorship theory. The Beinecke also describes the Bacon association as speculative. Yale News on the manuscript’s material history.
What would reopen it
Not direct Bacon authorship of the surviving object. That conflicts with the material chronology.
A weaker hypothesis could remain logically possible: a fifteenth-century maker copied or transformed an older text ultimately associated with Bacon or a Baconian tradition. That would require independent textual or historical evidence connecting the Voynich content to a traceable Bacon source. The old attribution itself is not enough.
Contradicted: Roger Bacon as physical author of the surviving codex.
Possible in the abstract but unsupported: an older Baconian source behind a later copy.
2. Wilfrid Voynich Forged the Manuscript
Why it was tempting
Voynich was a rare-book dealer. He acquired the manuscript in 1912. He hoped its decipherment and possible Bacon connection would increase its value. A mysterious object entering the modern market through a dealer naturally invites forgery suspicion.
What survived
It is always legitimate to investigate the incentives and documentary chain around an antiquities dealer. Provenance should never be accepted merely because the seller tells a compelling story.
What broke
The manuscript was known centuries before Wilfrid Voynich was born. A seventeenth-century correspondence trail connects Georg Baresch, Johannes Marcus Marci and Athanasius Kircher; Yale notes a 1639 Kircher reply as the earliest known documentary reference. The Marci letter accompanied the manuscript. Jacobus Horčický de Tepenec’s erased ownership mark belongs to the earlier Prague history. The parchment dates to the fifteenth century, and material analyses found inks, pigments and construction consistent with medieval manuscript production rather than modern manufacture.
These independent layers jointly defeat the simple claim that Voynich manufactured the codex in the early twentieth century. Beinecke’s manuscript overview and provenance.
What would reopen it
Only extraordinary evidence capable of explaining away the seventeenth-century documentation and the material chronology simultaneously. Repeating “he was a bookseller who benefited from the mystery” does not do that work.
Contradicted: a simple twentieth-century forgery manufactured by Wilfrid Voynich.
3. Rudolf II’s Roger Bacon Belief Is Verified Provenance
Why it was tempting
Marci’s letter is a real document, and it reports a specific story: Rudolf II bought the manuscript for 600 ducats and believed Bacon was the author. Later retellings often compress “Marci reports that he was told” into “Rudolf owned it and knew it was Bacon”.
What survived
The story is historically significant. It may preserve genuine information about the manuscript’s earlier imperial association. Rudolf’s court is a plausible environment for an unusual book to be collected, discussed and valued.
What broke
The letter is not a surviving imperial receipt, catalogue entry or first-person Rudolf statement. The Bacon claim is wrong for direct authorship. Therefore the reported story contains at least one demonstrably unreliable component. That does not make the entire Rudolf story false; it lowers the confidence with which each element should be repeated.
What would strengthen it
Independent imperial inventory evidence, correspondence, payment records or an earlier ownership mark directly connecting MS 408 to Rudolf would be far stronger than repetition of Marci’s report.
This is why our Broken Provenance Chain article distinguishes direct custody evidence from inherited stories.
Reported provenance ≠ independently verified provenance.
4. The Whole Manuscript Is One Plant → Preparation → Body → Astrology → Recipe Workflow
Why it was tempting
Few Voynich theories are more elegant.
The herbal pages identify plants. The fragment pages select useful parts. The vessels store or prepare those parts. Quire 13 shows treatment or bodily application. Zodiac pages supply timing. The starred entries record recipes.
The entire manuscript suddenly becomes one working medical machine.
Historical comparators make every individual link conceivable.
What survived
The manuscript genuinely changes representational scale. Some whole-plant and fragment relationships remain interesting. Medical manuscript culture genuinely could connect materia medica, celestial timing, bodily treatment, pharmacy and recipes. Vessels and segmented entries may well be functionally organised. The visual regimes do not look like unrelated random decoration.
What broke
The universal sequence did not generalise strongly enough across the whole manuscript. Too many links depended on assigning function from appearance before text independently supported it. Present folio order cannot safely be treated as original workflow. Visual section boundaries do not map neatly onto Currier regimes, hand variation or physical gatherings. Some supposedly sequential page families have strong internal structures that do not reduce naturally to one pipeline.
In other words, the pieces remain interesting while the single universal machine remains unearned.
What would reopen it
A repeated, predictive cross-regime mapping. For example, a robust rule linking a whole plant to a specific fragment, a fragment to a vessel class, that class to a Quire 13 operation and then to a starred entry—across many independent examples, including examples not used to formulate the rule.
Plausible local relationships survive. The universal pipeline remains unsupported.
5. One Distinctive Plant Feature Identifies the Species
Why it was tempting
A leaf resembles ivy. A root looks like ginger. A flower resembles a poppy. Once a familiar feature appears, the human visual system completes the object with astonishing confidence.
Then the adjacent text becomes irresistible as a candidate plant name.
What survived
Some Voynich plants may indeed preserve enough diagnostic morphology for identification or family-level classification. Plant images are not useless simply because many are ambiguous.
What broke
Single-feature matches are rarely unique. Medieval herbal copying can distort proportions, colours and parts. Composite-looking plants may reflect stylisation, copying error, schematic emphasis or genuinely unfamiliar species. Different modern readers routinely identify the same Voynich plant as different species because they weight different features.
A theory that can identify one plant as whichever species its proposed translation requires is not being constrained by the image.
What would strengthen an identification
Multiple independent diagnostic features; a historically plausible geographic range; comparison with relevant medieval image traditions; internal recurrence; label behaviour; and ideally predictive success on other plant pages.
One resemblance can suggest a species. A species claim needs a constellation of constraints.
6. Every Ornate Vessel Is a Pharmacy Jar
Why it was tempting
Late medieval and Renaissance pharmacies used visually distinctive storage vessels. Voynich contains tall and elaborate container-like forms beside plant fragments. The comparison is historically grounded.
What survived
A pharmaceutical comparison remains plausible. The page architecture looks more modular and inventory-like than the full-page herbal regime, and real apothecary culture provides an appropriate material control.
What broke
No exact vessel identification has turned the accompanying labels into independently confirmed contents. Some drawn structures may be decorative, symbolic or schematic rather than literal ceramic types. “Looks like an albarello” cannot do the work of proving ingredient, function and geography simultaneously.
What would strengthen it
A systematic correspondence among vessel morphology, known historical container classes, label families and plant-fragment classes—preferably one that predicts unseen examples.
7. The Zodiac Is the Master Control Layer for the Entire Book
Why it was tempting
Medieval medicine could involve celestial timing. Zodiac pages are among the manuscript’s clearest iconographic anchors. If the book concerns plants or treatments, an astrological timing layer could elegantly connect them.
What survived
Medical astrology is historically real. A relationship between the zodiac pages and some medical or calendrical function remains plausible. The recognisable signs offer useful controlled contexts for comparing unknown labels.
What broke
No manuscript-wide mapping has demonstrated that zodiac divisions systematically control herbal selection, Quire 13 states, vessel use or Quire 20 entries. The theory becomes circular when every date-like or repeated feature elsewhere is interpreted as confirmation simply because the zodiac theory predicts timing somewhere.
What would reopen it
Repeated correspondences between specific zodiac units and independently defined textual or visual classes elsewhere, with clear exclusions and predictive performance.
Zodiacal structure is real. Manuscript-wide control by that structure is not established.
8. The Rosettes Foldout Is a Specific City
Why it was tempting
Towers, walls, causeway-like connections and differentiated regions create a map-like impression. The huge foldout seems to invite geographic reading, and researchers have proposed cities, landscapes and regional networks.
What survived
The foldout genuinely encodes a relational spatial system. Nine major circular zones occupy a structured field linked by pathways or bands. Geometry, connectivity, enclosure and centre-periphery relations are real features worth modelling.
What broke
No proposed city has won by independently predicting the full nine-node topology, orientation, distinctive structures and textual placement. City theories often begin with one suggestive tower or landscape feature and then flex the remaining geometry to fit.
What would strengthen a geographic reading
A mapping specified before interpretation of individual details, with diagnostic landmarks, connectivity, orientation and scale relationships that distinguish the proposed geography from rival places and from non-geographic diagram types.
Geometry before geography.
9. Current Page Order Reveals the Original Workflow
Why it was tempting
A bound book trains us to read sequence as intention. Page A precedes page B, therefore A conceptually precedes B.
What survived
Some present local sequences may well preserve original order. Continuous bifolia, text continuity and consistent page families can support that.
What broke
The manuscript has missing leaves, unusual foldouts and a binding history. Quire reconstruction shows that physical order is an evidentiary question rather than a default assumption. A modern folio sequence cannot simply be converted into a medieval process diagram.
What would establish workflow
Repeated process transitions supported by physical continuity, cross-reference-like structure, consistent visual transformation or textual dependence—not mere adjacency.
Current order ≠ original order ≠ conceptual order ≠ process order.
10. The Missing Leaves Contained the Key
Why it was tempting
Voynich is incomplete. Once a theory encounters a missing explanation, the lost folios become a convenient place to store it.
The missing first page had the alphabet.
The lost foldout had the map key.
The removed recipe page explained the symbols.
What survived
Missing folios genuinely increase uncertainty. They may have contained structurally important material. A lost bifolium can interrupt text or image continuity.
What broke
Absence cannot be used as positive support for a specific theory. A hypothesis that survives only because every missing prediction is assigned to a missing leaf has made itself impossible to test.
What would make a lost-content inference legitimate
Physical reconstruction showing exactly where material is missing, combined with continuation cues on surviving leaves that constrain what kind of content likely occupied the gap.
Missing evidence can explain uncertainty. It cannot supply whatever evidence a theory lacks.
11. External Resemblance Is Provenance
Why it was tempting
A Voynich root resembles Carrara. A vessel looks Italian. A castle seems Lombard. A zodiac style feels French. A marginal form looks German.
Each match may be genuinely interesting.
What survived
Comparators can rank historical environments, identify iconographic families and reveal plausible production traditions. Regional clustering may eventually contribute to localisation.
What broke
No isolated visual resemblance establishes a transmission chain. Search bias can create false convergence: if we inspect enough manuscripts from one favoured region, we will accumulate many regional cousins while ignoring alternatives elsewhere.
What would strengthen localisation
Several independent regional indicators that are genuinely diagnostic and difficult to reproduce in competing regions: palaeography, material sourcing, linguistic annotation, iconographic convention, manuscript construction and documented transmission all pointing in the same direction.
Resemblance can rank possibilities. Provenance requires a chain.
12. A Later Readable Annotation Reveals the Main Language
Why it was tempting
The manuscript contains later writing that looks far more familiar than Voynichese: month names around zodiac diagrams, the Tepenec ownership mark and the difficult but partly recognisable-looking marginal material on folio 116v.
What survived
These additions are valuable evidence about reception, custody, readers and the environments through which the manuscript travelled.
What broke
The readable or partly readable forms belong to different hands or chronological layers. A later Romance month name cannot simply become the plaintext language of the primary script. A German-looking marginal phrase cannot silently reclassify every herbal page as German.
What would connect the layers
A repeated bilingual or explanatory relationship showing that a later annotation explicitly glosses a Voynich token or text passage in a reproducible way.
Later language is evidence about later hands until a bridge to the primary script is demonstrated.
The stratigraphic treatment is in Marginalia and the Later Hands.
What the First Twelve Defeats Already Teach
Notice the pattern.
Most failed theories did not begin from nothing.
They began from a real clue and then asked the clue to do too much.
- A historical attribution became authorship.
- A dealer’s incentive became fabrication.
- A reported ownership story became a verified custody document.
- A plausible multi-domain relationship became one universal workflow.
- A plant resemblance became species identity.
- A vessel resemblance became pharmaceutical function.
- Zodiac imagery became manuscript-wide control.
- Relational geometry became one named city.
- Present adjacency became original process order.
- Missing leaves became containers for missing proof.
- Regional resemblance became provenance.
- Later annotations became the language of the primary script.
This is the central disease of circular Voynich reasoning:
A true observation is promoted into a larger claim without paying the evidentiary cost of the promotion.
The next set of loops happens inside the writing itself.
13. One Visible Glyph Equals One Ordinary Alphabetic Letter
Why it was tempting
The script looks letter-like. It runs from left to right. Many recurring shapes have stable forms. Once EVA gives those forms familiar Latin-keyboard names, the eye begins to feel that the manuscript is one substitution table away from reading.
What survived
The script genuinely has recurring visual units. Treating those units as characters is a useful working model, and alphabet-like behaviour remains one serious possibility.
What broke
Segmentation is not settled. Some familiar Voynich forms may be composites, ligatures, modifiers or multi-part constructions. Conversely, what looks like one connected sign may contain several functional units. A one-visible-shape → one-plaintext-letter assumption also struggles to explain many of the manuscript’s positional restrictions without adding further machinery.
What would strengthen the simple-letter model
A stable glyph inventory whose members map consistently to an underlying phonological or alphabetic system across Currier regimes, scribal variation, labels, prose and line positions—with few special exceptions.
Visible sign ≠ proven plaintext letter.
14. Every Visible Space Is an Ordinary Word Boundary
Why it was tempting
The gaps are plainly there. Readers and transcription systems need a way to segment the text. Calling the chunks between gaps “words” is convenient and often analytically productive.
What survived
Space-delimited tokens are one of the strongest visible organisational features of Voynichese. They have repeatable length distributions, internal restrictions and families.
What broke
Nothing in the manuscript independently proves that each space delimits one lexical word in the modern linguistic sense. Spaces could delimit syllabic groups, abbreviation units, cipher blocks, formula components or another functional unit. Treating the visible token as a lexical word and then using “word behaviour” to prove the assumption is circular.
What would strengthen ordinary word segmentation
A decipherment or partial bilingual relation in which visible spaces repeatedly align with independently recoverable lexical boundaries, including difficult line-end and label contexts.
Stable token boundary ≠ demonstrated lexical word boundary.
15. q Has One Fixed Meaning
Why it was tempting
EVA q is one of the most conspicuously position-sensitive glyphs in the manuscript. It occurs overwhelmingly at visible token beginnings and is very often followed by EVA o. That is exactly the kind of regularity researchers hope will reveal an operator, prefix, article, numeral marker or grammatical feature.
What survived
q-series forms are genuinely constrained and deserve separate modelling. Their distribution is not what we would expect from a freely interchangeable ordinary alphabetic sign.
What broke
The positional regularity does not choose among prefix, ligature, cipher control, abbreviation, syllabic onset, modifier or generated-slot explanations. A fixed semantic gloss can often be made to fit selected tokens while failing to explain all q-bearing contexts.
What would establish a fixed role
A mechanism that predicts when q must occur, when it cannot occur, what changes when it is removed or replaced, and how the same rule behaves across Currier A/B, line positions and discourse types.
q has a strong distributional role. Its semantic role is not fixed by distribution alone.
16. Gallows Are Capitals, Numbers or Operators — Settled
Why it was tempting
Gallows are tall, conspicuous and often occur in visually privileged positions. Some cluster near paragraph or line starts. In known scripts, unusual initial forms can mark capitals, section signs, numerals, abbreviations or control symbols.
What survived
Gallows form a meaningful visual and positional family. Their architecture and distribution are worth modelling separately from short common glyphs.
What broke
No single capital, numeral or operator interpretation has yet explained the complete set of simple gallows, bench/pedestal combinations, internal token occurrences and cross-regime behaviour. The initial-position analogy is real; the semantic promotion remains unearned.
What would reopen a specific interpretation
A rule that accounts for the entire gallows family, not merely the most visually dramatic paragraph starts, and predicts their occurrence in unseen text.
17. Currier A and B Are Automatically Two Different Languages
Why it was tempting
Currier himself used language terminology, and the statistical contrast between A-like and B-like pages is strong enough to feel dialectal or linguistic.
What survived
Two broad textual regimes are one of the most durable findings in Voynich studies. Any serious theory should explain the distinction.
What broke
The observed distribution does not identify its hidden cause. Language, dialect, register, topic, scribe, production period, encoding mode and document function can all change surface frequencies. Some of those factors may act together.
What would identify two languages
A recoverable linguistic mapping in which the two regimes independently correspond to different grammatical or lexical systems, while ruling out simpler production and encoding causes.
Currier A/B = robust partition. Currier A/B ≠ two languages by definition.
18. Five Scribes Is a Settled Count
Why it was tempting
Lisa Fagin Davis’s 2020 palaeographic classification gave researchers a clear, usable five-hand model and immediately generated productive comparisons with Currier groups and manuscript sections.
What survived
Handwriting variation is real and analytically important. Treating scribal variation as a first-class variable was a major advance.
What changed
The 2026 critique by Torsten Timm challenged whether the variation is best partitioned into five discrete hands, arguing that a more continuous model may explain some of the same evidence. The correct 2026 public position is therefore not “there are five scribes” but “a five-hand model is influential and now explicitly contested”.
What would settle the count
Independent palaeographic analyses using clearly defined, replicable diagnostic features, ideally integrated with material, quire and production evidence rather than inferred from textual regime alone.
Real variation ≠ settled number of writers.
19. Find a Target Language, Then Make the Tokens Fit It
Why it was tempting
Voynich contains enough tokens that many languages can produce seductive local coincidences. If the researcher already suspects Latin, Italian, Hebrew, Nahuatl, Turkish, Romani, German, an Indic language or another candidate, a few tokens will almost inevitably resemble words after substitutions, deletions, abbreviations, anagrams, vowel insertion or flexible segmentation.
What survived
Candidate-language testing is legitimate. Historical geography, phonotactics, morphology and known scribal conventions can make some languages better candidates than others.
What broke
Flexible mappings can fit almost anything. If the same Voynich glyph may represent several sounds when needed, if vowels can be inserted freely, if token boundaries can be ignored selectively, if letter order can change, and if inconvenient words are treated as abbreviations or nulls after the fact, the candidate language is no longer making predictions. The researcher is editing the ciphertext until it resembles the answer.
What would make a language proposal serious
Publish the mapping rules before applying them broadly. Keep them stable. Translate long stretches, not isolated gems. Explain grammar as well as vocabulary. Predict unseen material. Account for failures rather than quietly changing the rule.
A language hypothesis should constrain the reading. The reading should not continuously redesign the language hypothesis.
20. Three Good Phrases Are a Decipherment
Why it was tempting
A readable phrase is emotionally powerful. If a short label can be turned into “red root”, or a zodiac string into a plausible month expression, the manuscript suddenly feels open.
What survived
A correct partial reading could begin with one phrase. Historically, decipherments sometimes do start from a small anchor.
What broke
Voynich’s flexibility makes local success cheap. A proposed rule that translates three hand-picked labels but cannot survive ordinary prose, Currier variation, repeated words, line effects and unseen folios has not deciphered the manuscript. It has produced three compatible readings.
What turns a reading into decipherment
Scale, consistency, grammar, prediction, reproducibility and resistance to cherry-picking. The solution must become harder to manipulate as more text is added.
Interesting reading ≠ manuscript-scale decipherment.
21. Language-Like Statistics Prove Ordinary Plaintext
Why it was tempting
Multiple studies have found long-range organisation, token clustering and other properties compatible with meaningful language. That is real evidence against the simplest idea of independent random characters.
What survived
Voynichese is highly structured. Its distributions deserve linguistic analysis. Simple shuffled-random baselines fail to capture important properties.
What broke
Structured cipher systems and rule-governed generators can reproduce language-like properties. The 2025 Naibbe work is particularly useful because it demonstrates that meaningful Latin or Italian plaintext transformed by a historically plausible hand cipher can produce several Voynich-like surface statistics. That keeps language underneath possible while defeating the inference from surface structure to ordinary visible plaintext.
Language-like structure ≠ ordinary plaintext.
22. Generated-Looking Structure Proves Meaninglessness
Why it was tempting
Near-neighbour word families, local repetition and low character-level uncertainty can be reproduced by simple generative rules. The surface sometimes looks as though a scribe transformed nearby tokens rather than independently choosing lexical words.
What survived
Local generation remains a serious mechanism family that any natural-language theory should be able to distinguish itself from. Some Voynich properties really are easy to reproduce with constrained generation.
What broke
Generated surface structure does not determine whether the input was meaningful. Encryption, abbreviation, codebooks and other transformations can generate highly artificial-looking outputs from meaningful sources. Conversely, a meaningless generator can imitate some language-like organisation. Surface generatedness does not settle semantic content.
Generated-looking ≠ meaningless.
23. One Entropy Number, Zipf Curve or Statistic Decides the Manuscript
Why it was tempting
A single scalar statistic feels wonderfully objective. If Voynich lies inside the natural-language range on metric X, perhaps language wins. If it lies outside, perhaps language loses.
What survived
Entropy, frequency rank, vocabulary growth, word length, network structure and intermittency are all useful constraints. They expose aspects of Voynich that visual intuition misses.
What broke
Different mechanisms can converge on the same scalar measure. Worse, transcription and segmentation choices can move the measurement itself. A theory that fits one statistic while failing line structure, local vocabulary or cross-regime behaviour has fitted a projection of the manuscript, not the manuscript.
What a statistic should do
Join a vector of independent constraints. The strongest model is the one that explains many difficult properties simultaneously with few arbitrary exceptions.
One metric can reject a model. One metric rarely identifies the historical mechanism.
24. My Model Produces Voynich-Like Statistics, Therefore It Made Voynich
Why it was tempting
Building a synthetic system that reproduces famous Voynich properties is genuinely impressive. It proves that the proposed mechanism family can reach part of the observed behaviour.
What survived
Constructive models are among the best ways to test claims of impossibility. They force vague theories to become executable and expose which properties follow from which rules.
What broke
A model tuned to reproduce known Voynich statistics may reproduce them precisely because those statistics were its design targets. Historical identity requires predictions outside the fitting set, a plausible fifteenth-century implementation, and manuscript-scale agreement with structures not used to build the model.
Simulation similarity ≠ historical identity.
25. One Whole-Plant ↔ Fragment Match Proves a General Cross-Section Dictionary
Why it was tempting
An apparent visual match between a full herbal plant and a later root or leaf fragment is exactly the sort of internal bridge we want. If accompanying labels also resemble one another, the discovery can feel decisive.
What survived
Internal whole↔part recurrence remains one of the more promising ways to constrain image and text together. It deserves systematic investigation.
What broke
A single match does not establish a manuscript-wide rule. Generic roots and leaves produce accidental similarities, and a flexible label comparison can find near-relatives almost anywhere in Voynichese.
What would establish the mapping
A pre-specified visual matching rule that succeeds across many eligible whole plants and fragments, paired with a textual relationship that also generalises and predicts held-back examples.
One bridge is a clue. A network of independent bridges is an architecture.
26. AI Found a Pattern, Therefore AI Translated Voynich
Why it is tempting
Modern machine learning can find clusters, similarities and statistical regularities that humans miss. Large language models can generate fluent translations from almost any input if prompted to search for one. Image models can rank visual similarities across huge collections. This creates the appearance of unprecedented decoding power.
What survives
AI and computation are genuinely useful for corpus search, anomaly detection, handwriting measurement, image retrieval, clustering, transcription comparison, candidate-model stress tests and the discovery of correlations too large for manual inspection.
What breaks
A model can always produce an interpretation if the task is “make this look meaningful”. Fluency is not evidence. An opaque classifier can distinguish Currier groups without knowing what Currier groups mean. A neural network can align a Voynich plant with a modern photograph without proving species identity. A language model can translate EVA into beautiful English while silently inventing the mapping.
The evidentiary burden does not decrease because the tool is sophisticated.
What would count as an AI-assisted breakthrough
Explicit reproducible rules or model outputs that make successful predictions on unseen manuscript material, survive independent replication, respect palaeography and codicology, and outperform plausible alternative explanations without post-hoc relabelling.
Pattern discovery is not translation. Fluent output is not decipherment.
27. “It Looks Italian/French/German” Is Enough to Locate the Manuscript
Why it was tempting
Humanistic scripts, zodiac month names, castle-like details, clothing, plant traditions and vessel forms all carry regional possibilities. Scholars rightly use style to localise manuscripts.
What survived
Regional art-historical and palaeographic analysis is a legitimate route. Individual features can change the probability of one production environment relative to another.
What broke
The fifteenth-century manuscript world was mobile. Artists copied older models. Books crossed borders. scribes moved. Intellectual content travelled through multiple languages. Search bias also matters: once a region is suspected, researchers preferentially find supporting comparators there.
What would locate it strongly
Convergence among independent diagnostic features—script, materials, quire practice, iconography, dialectal annotation, known workshop habits and documentary history—combined with active comparison against competing regions.
Regional fit ≠ regional proof.
28. “Hoax” Is One Simple Yes/No Hypothesis
Why it was tempting
The word hoax feels like the opposite of meaningful text. If the manuscript is not readable, perhaps somebody fabricated nonsense to fool a patron.
What survived
A fifteenth-century production whose visible text is partly or wholly non-semantic remains logically possible. Medieval people were capable of deception, play, pseudo-writing, magical display and invented notation. The old assumption that an elaborate medieval book must contain ordinary recoverable prose is not guaranteed.
What broke
“Hoax” often collapses incompatible hypotheses: a twentieth-century forgery, a fifteenth-century fake sold for money, meaningful encrypted text, mnemonic notation, ritual pseudo-writing and algorithmically generated filler. Evidence against one does not eliminate the others. The twentieth-century Voynich-forgery version is strongly contradicted; the broader possibility of medieval non-ordinary text remains open.
Also, non-semantic does not imply careless. A meaningless or partially meaningful system can still be highly structured.
“Hoax” is too broad to test until you specify who, when, how, why and what part of the object is supposedly deceptive.
29. One Partition Explains Everything
Why it was tempting
Researchers want the manuscript to snap into one clean architecture: visual sections = topics = scribes = Currier languages = production stages = original quires.
What survived
All of those partitions contain information. Their correlations are worth measuring.
What broke
The partitions do not align perfectly. A visual regime can contain more than one textual behaviour. A hand model can cut across subject type. Labels can form a functional class across sections. Physical gatherings do not reduce cleanly to modern section names.
Forcing one partition to become the master partition throws away the very mismatches that may reveal production history.
Voynich may be layered rather than neatly boxed.
30. A Statistical Keyword Means the Object Beside It
Why it was tempting
If a token disproportionately occurs in herbal text, calling it “plant”, “root” or a species name feels natural. If it clusters in zodiac pages, perhaps it means “star”, “month” or “degree”.
What survived
Distribution can tell us that a form carries information about section, local context or function. That is valuable.
What broke
The same distribution could reflect topic, grammar, scribe, encoding mode or document function. A herbal-page keyword might be an article used by one scribe rather than the word “plant”. Naming the token from the illustration and then using the illustration to confirm the name is circular.
Statistical association can locate a token before it can translate it.
31. The Same Star Shape Has the Same Meaning Everywhere
Why it was tempting
Stars occur around zodiac diagrams, in the hands of figures and beside Quire 20 entries. Repeated symbols beg for a universal dictionary.
What survived
The visual recurrence itself is worth mapping. It may reveal design continuity or a family of related concepts.
What broke
Modern symbols demonstrate why form alone is insufficient: an asterisk can mean multiplication, footnote, wildcard, correction or rating. Voynich stars occupy visibly different structural roles. A universal meaning must be demonstrated across those roles, not assumed from shape.
Same shape ≠ same function.
32. Reading EVA Aloud Gets Us Closer to the Language
Why it was tempting
EVA strings look pronounceable enough that our brains automatically vocalise them: daiin, chedy, qokeedy. Once a sound appears in the mind, it becomes easy to compare it with words from known languages.
What survived
Pronounceable aliases are useful memory aids for researchers.
What broke
EVA was not designed as a phonetic decipherment. Treating its Roman letters as sound values introduces accidental English/Latin-letter associations that are not manuscript evidence.
EVA spelling is a convenient label for visible shapes, not recovered pronunciation.
33. If No Medieval Manual Describes the Mechanism, the Mechanism Is Impossible
Why it was tempting
Historical plausibility matters. An elaborate cryptographic or notational proposal becomes less attractive if it requires technology, concepts or procedures unavailable to a fifteenth-century maker.
What survived
Any proposed mechanism must fit the material and intellectual capabilities of the period. A computer-dependent procedure with no hand-executable analogue is a poor historical candidate.
What breaks if we demand an exact surviving instruction manual
The survival record is incomplete. People can invent procedures without writing theoretical treatises about them. Workshop habits, personal ciphers, mnemonic devices and practical abbreviations may leave no separate explanatory document. Absence from surviving manuals therefore lowers historical support but does not automatically prove impossibility.
The correct standard
Ask whether the mechanism could be executed with fifteenth-century materials, time, numeracy and scribal skills, and whether comparable operations or conceptual ingredients existed. Historical possibility is a gate, not an identity certificate.
Not documented exactly ≠ historically impossible. Historically possible ≠ historically used here.
The Frozen Anti-Loop Ledger
At this point we can freeze the recurring claims into four public buckets. The bucket is not eternal; new evidence can move a claim. But repetition alone cannot.
Contradicted by current evidence
- Roger Bacon physically wrote the surviving fifteenth-century codex.
- Wilfrid Voynich manufactured the manuscript as a simple twentieth-century forgery.
- EVA letters are established sound values merely because they use familiar Roman letters.
Unsupported as established claims
- The Rosettes foldout is one specific named city.
- A single universal plant → preparation → body → astrology → recipe workflow explains the whole manuscript.
- The zodiac is the demonstrated master timing/control layer for every section.
- One modern plant identification yields a secure Voynich dictionary entry.
- Every ornate vessel is a known pharmaceutical jar with decoded contents.
- Current page sequence is the manuscript’s original conceptual or process sequence.
- The missing leaves contained the alphabet, key or explanation needed by a favourite theory.
- A regional visual resemblance proves production in that region.
- Later Romance/German/Latin-looking annotations identify the primary Voynich language.
- Currier A/B automatically means two natural languages or two decoded topics.
- Five scribes is an uncontested fact.
- q, gallows or star forms already have one demonstrated universal semantic value.
- Three plausible phrases constitute a decipherment.
- One entropy, Zipf or language-likeness result identifies the mechanism.
- A synthetic model that resembles Voynich is therefore the historical method used.
- An AI-generated fluent translation is evidence merely because it is coherent.
Still plausible but unproven families
- Meaningful natural language beneath an unusual script or abbreviation system.
- Meaningful text transformed by a cipher, code, nomenclator or verbose encoding.
- A mixed system combining language, abbreviation, labels and specialised notation.
- Rule-generated text, whether meaningful at an underlying level or partly/wholly non-semantic.
- A medical, pharmacological, astrological, balneological, mnemonic or mixed practical-knowledge function.
- Northern Italian, central European or other European production environments consistent with the material horizon.
- Multiple scribes or one/more scribes whose practices changed over time.
- Systematic whole-plant ↔ plant-fragment relationships in at least some cases.
- Cross-modal image–text relationships that could eventually constrain partial semantics.
Surviving observations that should not be reset to zero
- The parchment is medieval and belongs to an early-fifteenth-century radiocarbon horizon.
- The codex is physically engineered, incomplete and historically rebound/handled.
- The main writing is highly structured and internally constrained.
- Space-delimited tokens, line effects, paragraph effects and positional glyph behaviour are real.
- Currier A/B is a robust distributional distinction even though its cause is unknown.
- Handwriting variation is real even though the exact scribe count is contested.
- Word families, local vocabulary and long-range organisation are genuine features.
- The visual programme contains multiple coherent regimes.
- The Rosettes foldout has real relational geometry.
- Quire 13 has real topology and repeated image–text architecture.
- Labels form an important functional text population.
- Quire 20 is visibly segmented into many addressable records.
- Text and image interact materially on the page.
- Historical comparators show that several otherwise strange subject combinations were possible in medieval manuscript culture.
- No accepted manuscript-scale decipherment presently turns these constraints into one reproducible reading.
When Is a Dead Theory Allowed Back In?
Whenever genuinely new evidence changes the premise.
Not when somebody rediscovers the same resemblance.
Not when a prettier visualisation makes the old correlation feel stronger.
Not when a language model writes a more fluent paragraph around the same three token matches.
A theory deserves reopening when something changes its evidentiary position:
- a new documentary source;
- a new material result;
- a better manuscript image revealing an ambiguous glyph;
- a reproducible cross-modal relationship across many folios;
- a mechanism that predicts unseen data;
- a new comparator that is diagnostic rather than merely similar;
- a stable decipherment rule that scales without exceptions;
- an independent replication of a controversial classification;
- a newly discovered manuscript that supplies a genuine bridge.
Then reopen it gladly.
Good research is not loyal to defeat.
It is loyal to evidence.
Why This Anti-Loop Ledger Matters More Than Another Solution Claim
Imagine a researcher arriving at Voynich ten years from now.
They should not have to spend their first year rediscovering that EVA is not phonetic, that the five-scribe count is contested, that statistical keywords do not translate themselves, that a city-like Rosettes interpretation needs geometry, that later month names belong to a different layer, or that the plant-to-recipe master pipeline is more elegant than demonstrated.
They should begin from the surviving frontier.
That is what mature disciplines do.
Physics students do not independently rediscover every failed ether theory before studying relativity. Medical researchers do not restart from every discarded mechanism of infection. Cryptographers preserve known attacks so the next system is not judged by vulnerabilities already understood.
Voynich research needs the same memory.
The frontier begins where the old loop ends.
Part VIII Reading Map
- What We Actually Know
- The Broken Provenance Chain
- EVA, Transcription and the Segmentation Problem
- Currier A and B
- The Scribe Problem
- Word Families
- Why Voynichese Is So Predictable
- What a Real Voynich Decipherment Must Survive
After Removing the Loops, What Is Left?
A surprising amount.
This is the point at which scepticism often makes its own mistake. We remove unsupported plant names, city names, language names, author names, fixed glyph meanings and master workflows—and suddenly it can feel as though nothing remains.
That is false.
The object is real.
The structure is real.
The differences are real.
The constraints are real.
The losses are real.
The uncertainty is real.
And all of those together constitute knowledge.
Part IX — What Actually Survived gathers that knowledge into one map.
Part IX — What Actually Survived
After Part VIII, Voynich can feel smaller.
Roger Bacon is gone as the physical author. The easy Wilfrid Voynich forgery is gone. The perfect city on the Rosettes foldout is gone. The one-plant-one-word dictionary is gone. The universal medical pipeline has been demoted. EVA has lost its imaginary sounds. Currier A and B have stopped pretending to be two already identified languages. Five scribes have become a contested model rather than a sacred integer. Statistics have been returned to their proper job as constraints rather than verdicts.
Have we destroyed the mystery?
No.
We have removed some of the fog around it.
And this is where an important intellectual mistake usually occurs. People confuse the removal of unsupported stories with the removal of knowledge. If we cannot name the language, plant, author, city or cipher, they conclude that we know nothing.
That is not what the surviving evidence says.
Unknown meaning does not imply unknown structure.
We know the object is real.
We know the structure is real.
We know the differences are real.
We know some later stories are wrong.
We know some current hypotheses remain alive.
We know where uncertainty is concentrated.
That is not ignorance.
It is a map.
The Lowest-Common Surviving Model
If we deliberately remove every interpretation that the evidence does not force, what is the strongest compact description left?
The surviving Voynich codex is a genuine medieval, deliberately structured manuscript artefact with constrained, non-uniform writing and coordinated visual systems. Its writing varies systematically across position, page, region and document function. Its imagery is organised into several recurring regimes. The mechanism that produced the writing, the degree to which it encodes recoverable linguistic semantics, the precise function of its visual programmes and its exact place of production remain open.
That paragraph is intentionally unspectacular.
It is also extraordinarily difficult to make disappear.
A natural-language solution must fit it.
A cipher solution must fit it.
A generated-text theory must fit it.
A medical interpretation must fit it.
A hoax theory must fit it.
A future AI model must fit it.
None gets to erase inconvenient parts of the manuscript in order to become elegant.
1. What Survived Physically
The most secure Voynich knowledge begins where mystery has the least room to manoeuvre: the material object.
The codex is genuinely medieval
Radiocarbon analysis of four parchment samples places the sampled skins in the early fifteenth century, commonly summarised as 1404–1438. Material examination of inks and pigments is compatible with historical manuscript production. The pages are calfskin parchment, physically cut, folded, gathered, written, illustrated and painted.
This does not date every stroke to an exact year. It does decisively move the surviving object out of the modern-fabrication fantasy world.
The book was engineered as a codex
Bifolia were folded and nested into gatherings. Foldouts were deliberately created where ordinary page geometry was insufficient. Some of those foldouts required planning before they could function as continuous information surfaces.
This is not loose doodling assembled accidentally into a later book.
The surviving codex is incomplete
Missing leaves are not a theory. Physical stubs, interrupted gatherings and collation evidence demonstrate loss. Some losses involve single leaves; others may involve bifolia or larger structural disturbances.
That means we do not possess every page the maker or makers originally intended to survive together.
Current order deserves caution
Because the codex has been handled, rebound, foliated later and damaged, current sequence cannot be treated as a perfect transcript of original conceptual order. Some local relationships are physically secure. Others remain less certain.
This is not a small codicological footnote. It changes every theory that turns present page sequence into process sequence.
The earliest secure documentary history is much later than production
The seventeenth-century Baresch–Kircher–Marci correspondence gives us the first well-grounded documentary layer around the manuscript. The Tepenec ownership mark provides an earlier Prague-associated ownership clue. Later custody becomes much clearer through Voynich, Ethel Voynich, Anne Nill, H. P. Kraus and Yale.
What remains is a large earlier gap.
The manuscript can be medieval without our knowing where it spent its first two centuries.
The physical object is older than its secure documentary biography.
2. What Survived About the Writing
The text remains unread, but it does not remain undescribed.
This distinction is the centre of modern Voynich work.
There is a limited recurring sign system
The main text repeatedly uses a relatively compact family of visible forms. Researchers disagree about exactly how some complex forms should be segmented, so the number of underlying units remains model-dependent. But the script is not an unbounded sequence of arbitrary drawings.
Its repeated components are stable enough to transcribe, compare and count.
The visible spaces are structured
Voynichese is divided into repeated space-delimited tokens. Those tokens have stable distributions and internal restrictions. Whether they correspond one-for-one to ordinary lexical words remains unknown, but their segmentation is not visually random.
Glyphs have positional preferences
Some signs prefer token beginnings. Others prefer middles or endings. q-series forms are highly initial. Gallows have distinctive distributions. Certain combinations recur while other conceivable combinations are rare or absent.
This is one of the strongest reasons to reject the idea that the visible sequence is unconstrained doodling.
Lines matter
The beginning and end of a written line are not statistically neutral positions. Some forms concentrate near starts, others near ends, and paragraph openings add another layer of difference.
This means a serious mechanism must explain not only which tokens exist but why physical line position influences their distribution.
Labels and prose are not interchangeable
Short strings attached to stars, figures, roots, vessels and diagram components overlap with the wider script but occupy a different document function. Their lengths, locations and distributions make them a natural sub-corpus rather than mere fragments of prose.
Same script, different job.
3. What Survived About Textual Difference
Voynichese is not one homogeneous statistical soup.
Currier A/B survived
Whatever ultimately causes the distinction, large groups of pages differ systematically enough that Currier A and Currier B remain useful broad textual regimes. Later statistical work has not made that contrast disappear.
What has disappeared is the right to call them two languages as though that part were already known.
Local vocabulary survived
Some token families are local. Neighbouring folios can share forms that are rare elsewhere. Certain strings strongly prefer one section or textual regime.
That creates a vocabulary geography across the manuscript.
The cause may be topic, scribe, source, encoding mode, production episode or a mixture. The locality itself remains.
Word families survived
Voynich tokens form dense networks of near-neighbours. Tiny visible changes create related forms. Recurrent beginnings, middles and endings make the vocabulary look internally templated or morphologically organised.
Natural-language morphology can produce this.
So can other mechanisms.
The families survive even while their explanation remains open.
Long-range organisation survived
Information-theoretic and network studies have repeatedly found non-trivial organisation across the corpus. Tokens cluster by page regions. Co-occurrence structure exists. Simple shuffling destroys properties present in the manuscript.
The safest conclusion is not “therefore natural-language plaintext”.
It is:
The corpus contains structured long-range dependencies that a complete model must preserve.
4. What Survived About Predictability
Voynichese is unusually constrained under common transcription schemes.
That sentence should remain in every serious description of the text.
Given part of a token, the set of likely next characters can be narrow. Character combinations occupy preferred positions. Many theoretically possible sequences rarely or never appear. The surface system is therefore more predictable than a naïve random model and, in some character-level analyses, more predictable than many ordinary alphabetic texts.
The cause remains open.
- phonotactic constraint;
- morphological template;
- orthographic convention;
- abbreviation;
- syllabic structure;
- cipher expansion;
- glyph segmentation;
- local generation;
- line-position rules;
- or combinations of these.
Low uncertainty is not a solution.
It is a mechanical obligation.
5. What Survived About Scribal Production
Handwriting variation survived the debate over how many hands exist.
Lisa Fagin Davis’s five-hand model remains an influential and productive classification. Torsten Timm’s 2026 critique makes the exact five-scribe count less secure by arguing that some diagnostic variation may be continuous rather than categorical.
The correct synthesis is not to average them into “maybe three scribes”.
It is to preserve what both views require us to notice:
The writing changes in systematic palaeographic ways across the codex, and those changes belong inside any account of production.
Whether the cause is five writers, fewer writers, one evolving principal hand, or a more complicated workshop pattern remains open.
6. What Survived About the Manuscript’s Partitions
The manuscript has several kinds of geography at once.
- physical geography: quires, bifolia, recto/verso, foldouts;
- visual geography: plants, circles, zodiac, Quire 13, vessels/fragments, starred entries;
- textual geography: Currier A/B and smaller local clusters;
- scribal geography: handwriting variation;
- functional geography: labels, prose, circular text, short records;
- lexical geography: local token families.
All of them survive.
None perfectly explains all the others.
That mismatch is not a nuisance to be averaged away.
It may be one of the best clues to how the manuscript was assembled.
A simple one-variable history would make the partitions line up. If one scribe wrote one topic in one quire using one textual mode, every classification should collapse onto the same boundary.
They do not.
Therefore a more realistic production model may need several interacting variables.
7. What Survived Visually
The images are not merely decorative noise.
Several coherent visual regimes exist
Pages with large plant figures look different from zodiac-like circles. Quire 13 differs from both. Vessel-and-fragment pages have another grammar. Quire 20 largely abandons figurative illustration while preserving repeated star markers beside short entries.
Those differences are not inventions of modern section labels.
What remains uncertain is what functional nouns should be attached to them.
The herbal regime is genuinely plant-like
Large drawings use repeated botanical morphology: roots, stems, leaves, flowers and branching. Some may correspond to real plants or inherited herbal images. Exact species assignments remain difficult and often non-unique.
The zodiac anchors are unusually strong
Several central figures are recognisable as conventional zodiac signs. This ties those pages to a historical zodiac iconographic tradition even though the surrounding Voynichese remains undeciphered.
The Rosettes geometry is real
Nine major regions occupy a deliberate relational field on a complex foldout. Connections, boundaries, centres and pathways can be studied without naming a city or cosmology.
Quire 13 has topology
Human figures do not simply float in decorative water. Containers and tube-like structures connect into repeated architectures. Labels and prose interact with those architectures. Whatever the subject, relations were worth drawing.
The vessel pages change scale
Whole plants give way to smaller fragments, labels and differentiated container-like forms. That shift in representational scale is observable even if “pharmaceutical” remains a conventional label rather than a translation.
Quire 20 contains explicit record-like segmentation
Repeated marginal star forms divide large amounts of text into many short, addressable units. “Recipe” remains a hypothesis. Segmentation is a fact of page architecture.
8. What Survived Across Image and Text
The image programme and text programme were not produced in ignorance of one another.
Writing fits around illustrations. Labels attach to graphic units. Diagram text follows spatial structures. Foldouts coordinate image and writing across surfaces too large for a normal folio. Quire 13 places text inside a networked visual environment.
This means page geometry belongs to the evidence.
It also means a pure character stream is an incomplete representation of the manuscript.
Spatial attachment can constrain semantic role without translating it. A string beside one object has a narrower range of plausible functions than unrestricted prose. A repeated image–label relation is stronger than a one-off association.
Some whole-plant↔fragment correspondences remain promising enough to deserve systematic investigation.
The universal cross-section pipeline does not survive at the same confidence.
Internal recurrence survived. Universal semantics did not.
9. What Survived Historically
The comparator work did not give us Voynich’s parent manuscript.
It gave us something more durable: a historically grounded possibility space.
From Carrara and related herbals we know that medicinal plant knowledge could be translated, copied, visually transformed and transmitted across languages.
From Egerton MS 747 we know that a medical codex could combine herbal material, a lunar calendar, an antidotary, doses, substitutions, weights, measures and synonyms.
From bath traditions we know that therapeutic waters could generate serious medical texts and illustrated manuscript programmes.
From surviving pharmacy material culture we know that differentiated storage vessels could belong to practical medical environments.
From known copying chains we know that images can drift while lineage remains real.
Therefore Voynich’s mixed visual world is not historically absurd.
But none of these comparators locates the manuscript by itself.
Historical coherence survived. Exact provenance did not emerge from resemblance.
10. What Survived About Language
The natural-language hypothesis is alive.
That is not the same as saying natural language has won.
Long-range organisation, token families, contextual vocabulary, line and paragraph structure, labels and other properties make linguistic analysis productive. The possibility that meaningful language exists beneath an unusual orthography, abbreviation system or transform remains serious.
What has not survived is the easy equation:
language-like statistics = identified natural language.
No language has produced a stable, accepted manuscript-scale reading under explicit rules.
Therefore natural language remains a live mechanism family, not a solved identity.
11. What Survived About Ciphering
The cipher hypothesis is alive too.
Simple monoalphabetic substitution is a poor description of many Voynich properties. But historically plausible ciphering includes much more than one-to-one substitution: homophones, nomenclators, nulls, abbreviation, verbose expansion, context-dependent rules and mixed systems.
The 2025 Naibbe demonstration matters because it shows that meaningful Latin or Italian can be transformed by a hand-executable verbose homophonic system into output sharing several famous Voynich-like statistical characteristics.
That does not identify the Voynich cipher.
It prevents some surface statistics from being used as blanket proofs that meaningful ciphertext is impossible.
Cipher remains possible. A specific historical cipher remains unproved.
12. What Survived About Generated Text
Rule-based generation remains possible.
Constrained algorithms can reproduce near-neighbour word families, repeated local forms, strong character ordering and long-tailed token frequencies. That means some Voynich properties are compatible with a generative procedure.
What does not follow is that the text is therefore meaningless.
A meaningful encoding can generate artificial-looking output. A mnemonic system can generate formulaic strings. A cipher can create families. A copyist can transform meaningful source material through local rules.
The surviving question is not simply:
Was Voynich generated?
It is:
If generation occurred, what was being generated from what, under which constraints, and with how much recoverable meaning?
13. What Survived About “Hoax”
The word remains too blunt.
A simple twentieth-century fabrication by Wilfrid Voynich conflicts with material and documentary evidence.
A fifteenth-century production containing partly or wholly non-semantic writing is a different hypothesis and is not ruled out merely because the object is genuinely medieval.
A private mnemonic system is different again.
A meaningful encoded manual is different again.
A ritual or display script whose function depends on appearance rather than propositional reading is different again.
Therefore the surviving lesson is methodological:
Never ask “Is Voynich a hoax?” until “hoax” has been decomposed into a testable historical mechanism.
14. What We Still Do Not Know
A mature knowledge map needs white space.
These are not embarrassing omissions. They are the actual open variables.
- Author: unknown.
- Commissioner or patron: unknown.
- Exact place of production: unknown.
- Exact date of writing: not fixed to a specific year.
- Underlying language or languages: unidentified.
- Whether the visible text directly encodes natural language: unresolved.
- Whether ciphering, code, abbreviation, generated structure or a mixed mechanism is involved: unresolved.
- Exact phonetic values of the primary glyphs: unknown.
- Whether visible spaces are ordinary lexical word boundaries: unresolved.
- Exact number of scribes: disputed.
- Exact original folio order: incompletely recoverable.
- Exact function of each visual regime: not decoded.
- Exact identities of most plant figures: uncertain.
- Exact referent of the Rosettes foldout: unknown.
- Exact meaning of Quire 13’s bodies, pools and tubes: unknown.
- Exact function of the vessel forms: unknown.
- Whether Quire 20 entries are recipes: plausible conventional description, not demonstrated semantics.
- Complete provenance before the seventeenth century: unknown.
- Plaintext: not recovered to scholarly consensus.
- Accepted decipherment: none.
That list is long.
It is shorter than “we know nothing”.
15. The Confidence Ladder Applied to Voynich
Now we can populate the evidence ladder introduced at the beginning of this master.
KNOWN
- A physical parchment codex survives as Yale Beinecke MS 408.
- Its sampled parchment falls within an early-fifteenth-century radiocarbon horizon.
- The surviving object contains writing, illustrations, coloured material, foldouts and later annotations.
- Leaves are missing.
- The codex has a seventeenth-century documentary history involving Baresch, Marci and Kircher.
- An ownership mark associated with Jacobus Horčický de Tepenec survives, though altered/erased.
- Wilfrid Voynich acquired the manuscript in 1912.
- The manuscript eventually entered Yale’s Beinecke Library in 1969.
- The main script is not currently identified with an accepted plaintext reading.
STRONGLY SUPPORTED
- The main writing is rule-constrained rather than independently random marks.
- The script contains stable recurring visual units, though exact segmentation is partly model-dependent.
- Visible tokens have strong internal positional restrictions.
- Line and paragraph position affect textual distribution.
- Currier A/B captures a genuine broad textual distinction.
- Vocabulary is locally structured across the codex.
- Handwriting varies systematically, even though the best discrete scribe count remains contested.
- The visual programme contains several internally coherent regimes.
- Text and illustration were coordinated on many pages.
- The Rosettes foldout deliberately represents a network of differentiated regions.
- Quire 13 deliberately represents repeated relationships among bodies, containers and connected forms.
- Quire 20 deliberately segments text into many star-marked or star-associated units.
PLAUSIBLE
- At least some of the writing ultimately carries meaningful linguistic content.
- The visible script represents language through unusual orthography, abbreviation, code, cipher or a mixed transformation rather than ordinary alphabetic spelling.
- At least some visual regimes belong to practical natural-philosophical, medical, pharmacological, calendrical or astrological knowledge.
- More than one production hand or changing scribal practice contributed to the surviving codex.
- Some large herbal plants and later fragments are intentionally related.
- Some labels encode stable object-related categories even if they are not straightforward names.
POSSIBLE
- A natural language under heavy abbreviation or transformed orthography.
- A verbose/homophonic or mixed cipher family.
- A constructed or mnemonic notation.
- A partly generated system.
- A partly or wholly non-semantic medieval textual production.
- Northern Italian production.
- Central European production.
- Another European production environment consistent with the material and stylistic evidence.
- A medical–astrological or broader practical-knowledge compilation.
- A production history involving several source exemplars.
UNSUPPORTED AS ESTABLISHED CLAIMS
- Padua, Milan, Prague or any other specific city as proven production location.
- One exact external herbal as the demonstrated source of Voynich’s plant programme.
- A manuscript-wide dictionary derived from plant pictures.
- One named city or landscape as the proven Rosettes referent.
- A universal plant→preparation→body→astrology→recipe workflow.
- The zodiac as a demonstrated control layer for all sections.
- Five scribes as an uncontested fact.
- Currier A/B as two identified languages.
- q, gallows, stars or other recurring forms as possessing one demonstrated universal semantic value.
- A fluent AI translation as a decipherment.
- Any existing proposed solution that cannot scale reproducibly across the manuscript.
CONTRADICTED
- Roger Bacon physically wrote the surviving fifteenth-century codex.
- Wilfrid Voynich manufactured the surviving object as a simple twentieth-century forgery.
- EVA’s Roman letters are known phonetic values merely by virtue of the transcription notation.
These categories are not commandments.
They are a current knowledge state.
New evidence may move an item upward or downward. But the movement should be caused by evidence, not repetition.
16. The Minimum Theory We Should Accept From Now On
This master can now impose a minimum standard on every future explanation.
A serious theory is no longer allowed to explain only the feature that made its author notice the manuscript.
It must preserve the facts that survived everybody else.
- It must fit the early-fifteenth-century material horizon.
- It must be compatible with actual codex construction.
- It must tolerate missing leaves and uncertain order without using them as escape hatches.
- It must explain a constrained recurring script.
- It must explain visible token structure.
- It must explain positional glyph behaviour.
- It must explain line and paragraph effects.
- It must explain Currier A/B or explicitly show why that distinction emerges secondarily.
- It must explain local vocabulary and near-neighbour word families.
- It must account for label/prose differences.
- It must accommodate handwriting variation.
- It must handle the multiple visual regimes without assigning meanings solely from appearance.
- It must explain why text and image are integrated spatially.
- It must be historically executable if it proposes a production mechanism.
- It must survive comparison with alternative mechanisms that produce similar statistics.
A theory that translates one herbal page beautifully and ignores the rest has not met the minimum.
A theory that reproduces entropy but cannot explain the labels has not met the minimum.
A theory that identifies the Rosettes but cannot tell us why the writing has Currier regimes has not met the minimum.
A theory that reads every word only by changing its rules has not met the minimum.
The next Voynich theory does not inherit a blank manuscript. It inherits a constraint set.
17. Why “We Still Cannot Read It” Is Not the Final Sentence
Imagine archaeology worked like popular Voynich discussion.
We would discover a ruined city, fail to identify the king, and announce that nothing about the city is known.
That would be absurd.
We could still measure the walls. Date timber. Reconstruct streets. Analyse diet. Identify trade goods. Map water systems. Distinguish construction phases. Infer population changes. Study burial practice.
The unknown king would remain important.
It would not erase the city.
Voynich is similar.
Translation is a major missing layer.
It is not the only possible knowledge layer.
We can reconstruct constraints around the writing long before we can assign semantic values to it.
And every genuine constraint makes the space of possible solutions smaller.
18. The Shape of the Remaining Mystery
At the beginning of this article, the Voynich mystery looked enormous because everything was mixed together.
Who wrote it?
What language?
What plants?
What city?
What cipher?
Why the women?
Why the jars?
After nine parts, the mystery has acquired a different shape.
We are no longer staring at one dark wall.
We have several narrower frontier questions:
- What underlying unit system best explains the visible glyph inventory?
- What mechanism produces the very strong positional restrictions?
- Why do line boundaries matter?
- What hidden variable generates Currier A/B?
- How much scribal variation is categorical and how much continuous?
- What creates dense word-family neighbourhoods?
- Why does vocabulary localise?
- Which image–text alignments generalise beyond chance?
- Can whole-plant↔fragment relationships be demonstrated predictively?
- Can visual geometry be represented computationally without flattening it?
- Can a historically plausible mechanism reproduce multiple text properties simultaneously?
- Can any mechanism yield stable recoverable semantics at manuscript scale?
- Can material or documentary evidence narrow the production region?
- Can an unknown comparator manuscript bridge the current provenance gap?
Those are much better mysteries.
They are narrower.
They are testable.
And they tell us where to spend the next hour rather than where to spend the next fantasy.
DON’T GO IN CIRCLES: The Surviving-Knowledge Edition
“If nobody can translate it, nothing has been learned.”
False. Physical, palaeographic, statistical, codicological, visual and distributional knowledge has accumulated substantially.
“If a theory has not been disproved, it is still equally likely.”
False. Possible is not the same as well-supported. Evidence can rank hypotheses without eliminating every alternative.
“If two explanations fit one statistic, statistics are useless.”
False. The statistic still excludes mechanisms unable to reproduce it. The next step is to combine independent constraints.
“Uncertainty means we should not say anything.”
False. Uncertainty should determine the strength of the sentence, not whether evidence can be discussed.
“A cautious conclusion is weaker research.”
False. A conclusion that survives future evidence is stronger than a spectacular claim built by spending uncertainty as certainty.
Part IX Reading Map
- What We Actually Know
- The Manuscript Before the Mystery
- The Missing Leaves and Quire Reconstruction
- What the Writing Does Before We Know What It Says
- Currier A and B
- The Scribe Problem
- Word Families
- Why Voynichese Is So Predictable
- What the Pictures Can and Cannot Tell Us
- What a Real Voynich Decipherment Must Survive
One Part Left
We have done the archaeology of certainty.
We know what is solid.
We know what is promising.
We know what is still merely possible.
We know which old loops no longer deserve to reset the conversation.
The final task is not another theory.
It is a research contract.
What should a future solution be required to do?
What can computation and AI genuinely contribute?
What data should be built next?
Which questions are now mature enough to answer?
And when somebody announces next year that the Voynich Manuscript has finally been solved, what should you ask before believing them?
Part X — Where Voynich Research Should Go Next answers those questions.
Part X — Where Voynich Research Should Go Next
The final task is not to propose another solution.
We have enough solutions.
What we need now is a better definition of what solving the Voynich Manuscript would actually require.
That may sound less exciting than announcing a language, a city or a cipher. It is also the work that makes a future announcement worth believing.
For nine parts, we have been narrowing.
We separated parchment from writing date.
We separated pictures from the names we give them.
We separated EVA from sound.
We separated Currier clusters from decoded languages.
We separated statistical structure from semantics.
We separated comparators from sources.
We separated plausible mechanisms from demonstrated historical mechanisms.
We separated useful failure from forgotten failure.
Now those separations become requirements.
The next Voynich solution does not begin with a blank manuscript. It inherits every constraint that survived the last century of work.
That is the research contract.
What Does “Solved” Mean?
The word solved is used far too cheaply around Voynich.
A plant identification is not a solution.
A plausible translation of three labels is not a solution.
A statistical classifier that separates Currier A from Currier B is not a solution.
A cipher that produces Voynich-like output is not a solution.
An AI model that generates fluent English from EVA is not a solution.
A map that makes the Rosettes look like northern Italy is not a solution.
All of those could become pieces of a solution.
But a solved manuscript requires the pieces to stop being independently adjustable.
The rules must harden.
The solution must become less free as the corpus becomes larger.
A genuine decipherment should reduce interpretive freedom rather than manufacture more of it.
A Partial Breakthrough Is Allowed to Be Partial
One reason solution claims become inflated is that researchers fear a partial result sounds unimpressive.
It should not.
If somebody proves that a particular label family marks zodiac degrees, that would be a major breakthrough even if 95 per cent of Voynichese remains unread.
If somebody demonstrates that one glyph pair is a scribal abbreviation rather than two independent characters, that may transform entropy estimates without yielding a sentence.
If a newly discovered manuscript demonstrates that one Quire 13 diagram derives from a known therapeutic-water tradition, that would be historically important without deciphering the text.
If material analysis localises a subset of pigments or parchment-production practices to a narrower region, provenance can advance without semantics.
Good research should be able to say:
We solved this relationship. We have not solved the manuscript.
That sentence should be celebrated, not treated as a failure of ambition.
The Sixteen Gates a Real Voynich Decipherment Must Survive
The following gates are deliberately demanding.
They have to be.
Voynich is large enough, repetitive enough and ambiguous enough that a weak method can manufacture local success almost indefinitely.
Gate 1 — State the representation rules exactly
What is a glyph?
Which visible forms are variants of the same unit?
Which forms are compounds?
How are spaces handled?
How are uncertain readings handled?
If the proposed solution silently changes glyph segmentation whenever a difficult word appears, everything downstream becomes unfalsifiable.
Gate 2 — Give a stable transformation or reading rule
If EVA q is interpreted one way on an herbal page and another way on a zodiac page, explain the rule that chooses between them before reading the examples.
If vowels are omitted, specify where and why.
If signs are nulls, specify the null rule.
If word boundaries move, specify the boundary mechanism.
A mapping is not stable if every exception creates a new mapping.
Gate 3 — Scale beyond showcase passages
A decipherment should work on ordinary pages, not only famous labels or hand-selected folios.
It should work where the pictures do not suggest the answer.
It should survive boring text.
It should survive repeated formulae.
It should survive pages chosen by somebody else.
A method that reads only the passages used to invent it has demonstrated compatibility, not generality.
Gate 4 — Predict unseen material
This is the gate most solution claims should fear.
Freeze the rules.
Then apply them to folios, lines or labels not used to construct the model.
The result should be better than a post-hoc interpretation created after seeing the answer context.
If a theory predicts that a particular token component marks one visual class, test it on all eligible instances, including those the researcher has not previously inspected.
Prediction is where resemblance becomes mechanism.
Gate 5 — Explain Currier A and Currier B
A proposed language or cipher cannot act as though Currier never happened.
Why do the broad textual regimes differ?
Different language?
Different dialect?
Different scribe?
Different encoding parameter?
Different genre?
The solution need not preserve Currier’s exact labels forever, but it must explain why the distributions that motivated them exist.
Gate 6 — Explain scribal and palaeographic variation
If five hands are real, why does the decipherment behave across them?
If one evolving hand is a better model, how does the textual mechanism change over time?
A solution should not confuse a spelling preference with a new word or a handwriting variant with a new cipher symbol without evidence.
The palaeography and the semantics have to coexist.
Gate 7 — Explain labels and running prose together
If labels are names, show how the same textual system produces prose.
If labels use a nomenclator or compressed form, specify the transition.
If labels and prose use different grammars, show why their shared token inventory still behaves as observed.
A proposed solution that works beautifully on labels and collapses on paragraphs has solved a subproblem at best.
Gate 8 — Explain line and paragraph effects
Why do token distributions change at line beginnings and endings?
Why do some tall forms favour privileged positions?
Why can first lines differ from interior lines?
Ordinary language translated into Voynichese by a proposed rule must somehow acquire those visible positional effects. If the rule has no mechanism for them, it has not explained the actual ciphertext—or whatever the surface system is.
Gate 9 — Explain word families and local vocabulary
Why do so many tokens have near-neighbours?
Why do related forms cluster?
Why are some tokens highly local?
Natural-language morphology can answer part of this. A local generator can answer part. A cipher can answer part. A good solution has to reproduce the actual distribution of similarities and absences rather than merely gesture at one mechanism family.
Gate 10 — Explain predictability, not just translate through it
The low conditional uncertainty of the visible character stream under common transcriptions is not optional background.
If the plaintext is an ordinary European language, what transformation makes the surface much more constrained?
If the script is syllabic or abbreviated, which rules generate the restriction?
If it is generated text, why does the generator produce the observed long-range organisation?
Translation and mechanics must meet.
Gate 11 — Respect the visual regimes without reading the pictures backward
If the translation says an herbal paragraph describes a plant, good.
Now test whether the same rules produce plant-relevant language on other herbal pages without using the images to choose meanings.
If a zodiac label translates as a date, test the other zodiac labels.
If Quire 20 produces recipes, identify the grammar of procedures, ingredients or quantities and test whether it behaves differently from descriptive prose.
The pictures may validate a translation.
They should not be allowed to manufacture it.
Gate 12 — Explain image and text jointly
The text was not written in empty space.
Labels attach to figures.
Prose wraps around plants.
Circular writing follows diagrams.
Foldouts preserve spatial relations.
A complete account should eventually tell us what those layout relationships are doing. Even if translation begins from the text alone, the interpretation should predict sensible cross-modal relationships rather than merely tolerate the pictures.
Gate 13 — Respect codicology and uncertain order
A proposed narrative should not require a page to follow another merely because the modern binding places it there.
If conceptual order matters, demonstrate it through physical continuity, text continuity, cross-reference, bifolium relation or another independent signal.
Missing leaves must create uncertainty, not convenient invisible evidence.
Gate 14 — Be historically executable
If the theory requires a transformation mechanism, could a fifteenth-century person actually execute it?
How much arithmetic?
What tools?
What tables?
How much memory?
How much time per line?
Could multiple scribes learn the convention?
A mathematically beautiful algorithm whose historical implementation requires a laptop is an excellent modern model and a weak fifteenth-century production hypothesis unless a hand-executable equivalent is demonstrated.
Gate 15 — Be reproducible and falsifiable
Another researcher should be able to take the same transcription, the same rules and the same folio and obtain substantially the same result.
The theory should also say what would make it wrong.
If every failure can be explained as a scribal error, null, abbreviation, homophone, special case or lost page, the model has no point of contact with defeat.
A theory that cannot lose cannot win.
Gate 16 — Beat the alternatives, not merely fit the manuscript
This is the hardest gate.
Voynich is flexible enough that several mechanisms can fit a chosen subset of properties.
A natural-language model may fit vocabulary clustering.
A generator may fit word families.
A verbose cipher may fit character restrictions.
A manuscript-function model may fit image sections.
The winning explanation must do more with less arbitrary machinery.
It should explain multiple independent constraints simultaneously and leave fewer unexplained residues than plausible rivals.
Fit is necessary. Discrimination is what turns fit into explanation.
A Seventeenth Gate: Meaning Has to Behave Like Meaning
If the claim is that Voynich contains language, the recovered output should eventually do what language does.
It should have syntax or another coherent combinatorial system.
Words or units should preserve meanings across contexts unless a principled polysemy rule explains change.
Pronouns, function words, inflections, particles, numbers or equivalent structural classes should behave consistently if the proposed language requires them.
Repeated passages should not turn into unrelated English because the picture changed.
Grammar should predict.
Semantic fields should recur.
Errors should look like plausible scribal errors rather than rescue clauses invented after the fact.
A meaningful text does not merely emit plausible sentences.
It constrains which sentences are possible.
A Real Decipherment Should Become Boring
This may be the strangest prediction in the entire article.
At first, decipherment would be thrilling.
Then it should become boring.
The fiftieth page should not require a miracle.
The hundredth repeated token should not require a new semantic improvisation.
The same grammatical rule should keep working.
The same cipher operation should keep producing valid output.
The same abbreviation should keep abbreviating the same kind of thing.
Scientists should begin arguing about what the recovered text tells us about fifteenth-century knowledge rather than arguing about whether the reading rules change on every line.
A mature decipherment replaces mystery with routine.
What AI Is Actually Good For
Artificial intelligence changes the scale of Voynich research.
It does not change the definition of evidence.
This distinction is going to become more important, not less.
A modern system can inspect thousands of manuscript images, compare millions of token contexts, cluster hand shapes, generate synthetic cipher corpora and search digitised libraries far faster than one human scholar.
Those are extraordinary capabilities.
The useful question is not “Can AI solve Voynich?”
It is “Which parts of the evidence problem become cheaper, broader or more reproducible when computation is used well?”
AI can search the corpus exhaustively
A human can notice that two tokens look similar.
A computer can rank every token against every other token under several similarity metrics, then tell us where each family occurs.
This is ideal for testing word-family claims, local vocabulary, repeated labels and cross-section recurrence.
AI can find anomalies
Rare glyph combinations, unusual line starts, abnormal pages, transitional Currier behaviour and outlier labels can be surfaced systematically.
Anomalies are valuable because they are places where a simple theory is most likely to break.
AI can compare competing transcriptions
Instead of treating one transcription as ground truth, models can propagate uncertainty across multiple expert readings.
If a statistical conclusion survives several segmentation schemes and transcription variants, it is stronger.
If it disappears when one ambiguous glyph is split differently, we learn that the conclusion was representation-sensitive.
AI can compare handwriting at scale
Digital palaeography can measure stroke shape, curvature, spacing, proportions and context across thousands of instances. This will not make human palaeographers obsolete. It can expose whether a proposed hand boundary is sharp, gradual or entangled with text type.
The 2020 five-hand model and 2026 critique make this an especially fertile area for independent replication.
AI can search manuscript images for comparators
Digitised libraries contain millions of folios no single researcher can inspect manually.
Image-retrieval systems can search for root structures, diagram layouts, vessel silhouettes, bathing configurations, star-marker forms and page compositions across collections.
But the output should be treated as candidate retrieval, not historical proof.
The machine can say “these look similar”.
Scholarship still has to ask whether the similarity is diagnostic, chronologically possible and historically connected.
AI can preserve geometry
Computer vision can represent where text blocks, labels, figures and lines sit on the page. Graph models can encode Rosettes connections or Quire 13 topology instead of flattening them into one linear string.
This may be one of the most important future advances because the manuscript is spatial by design.
AI can build counterexamples
If somebody claims “no meaningful encoded text could ever have property X”, computational modelling can attempt to construct one.
If a counterexample exists, the impossibility claim falls.
This is one reason synthetic corpora and explicit cipher models are valuable even when they are not historical identifications.
AI can stress-test candidate solutions
Once a proposed mapping is published, computation can apply it to the entire corpus rather than the examples selected by its author.
How many tokens become readable?
How many exceptions are required?
Do repeated tokens receive repeated meanings?
Does grammar remain stable?
Do the rules behave differently only when the pictures would otherwise contradict the reading?
This turns public decipherment claims into testable artefacts.
What AI Is Bad At
AI is especially dangerous where Voynich is already dangerous.
It is good at finding a plausible story.
Voynich contains enough ambiguity to reward plausible stories.
The combination can be spectacularly misleading.
Fluent translation without a fixed mapping
A language model can turn arbitrary input into fluent prose because producing coherent prose is one of its strengths.
If the mapping from Voynich signs to that prose is not explicit, stable and independently reproducible, fluency tells us almost nothing about the manuscript.
Pattern naming
An AI system may correctly detect that a cluster exists and then incorrectly name the cluster “herbal ingredients” because it knows the surrounding pictures show plants.
Discovery and naming must remain separate.
Opaque confidence scores
“The model is 87 per cent confident this means water” is not evidence if nobody can inspect what the 87 per cent refers to, what alternatives were compared, how the training data influenced the answer or whether the score is calibrated.
Comparator contamination
A model trained on thousands of images can retrieve objects that resemble Voynich—but if the search or interpretation is already seeded with “Italian herbal”, the output may simply explore the semantic neighbourhood of the seed.
Search breadth must be designed deliberately.
Retrospective explanation
AI is excellent at explaining why an observed result makes sense after seeing it.
Voynich requires models that state what should happen before seeing the case.
AI can accelerate the search. It cannot lower the evidentiary bar.
Or, put more simply:
A model is a microscope, not a witness.
The Dataset Voynich Actually Needs
Most digital Voynich work still inherits a hidden assumption from ordinary text processing: that the manuscript can be represented adequately as a sequence of characters.
It cannot.
A future corpus should treat each folio as a structured information surface.
1. Preserve coordinates
Every transcribed line, token and label should retain page coordinates. A label beside a root should not become indistinguishable from a label beside a star merely because both are stored in a text file.
2. Preserve orientation
Circular and radial writing should retain direction and orientation. A line read around a ring should not silently become equivalent to ordinary left-to-right prose.
3. Preserve document function
Label, paragraph, circular inscription, marginal annotation, quire number, folio number and star-marked entry should be tagged separately.
4. Preserve attachment
When a string is visually associated with a plant part, star, figure, vessel, tube or diagram region, that relation should be represented explicitly—with uncertainty where attachment is ambiguous.
5. Preserve multiple transcriptions
Do not force one glyph reading to become ground truth when qualified transcribers disagree. Store alternate readings and confidence.
Then ask whether conclusions survive the transcription ensemble.
6. Preserve segmentation alternatives
A bench form might be one unit or a composite. A connected shape might be a ligature. The corpus should allow competing segmentation models rather than hard-code one ontology forever.
7. Preserve codicology
Each token should know its folio, recto/verso, bifolium, quire, foldout and known physical relationships. Missing leaves and uncertain reconstruction should be encoded as uncertainty, not forgotten.
8. Preserve hand hypotheses
Hand assignments should be represented as models with provenance and confidence, not as an invisible fixed truth. Five-hand, alternative and continuous-feature representations should be comparable.
9. Preserve visual features
Plants should have computable descriptors for leaf arrangement, root topology, flowers and branching. Quire 13 should preserve figure–container–tube relations. Rosettes should preserve graph structure. Stars and vessels should have morphological descriptors.
10. Preserve uncertainty
“Unknown” should be a first-class value.
So should “ambiguous”, “damaged”, “probable”, “later hand”, “reading order uncertain” and “attachment uncertain”.
A dataset that converts uncertainty into forced categorical labels creates fake certainty that later models will learn as fact.
We Need a Codicological Graph, Not Just a Page List
Imagine the manuscript represented as a graph.
A folio connects physically to its conjugate leaf.
Both belong to a quire.
A foldout contains spatial regions.
A text block belongs to one region.
A label attaches to a figure.
A figure connects to a tube.
A token belongs to a line, paragraph, Currier classification and proposed hand.
A visual motif recurs on another folio.
Now the corpus can ask cross-scale questions without pretending every relationship is linear.
This is the kind of representation Voynich has been asking for all along.
The physical codex is already a graph.
Our data should stop flattening it.
We Need a Comparator Database Built Around Known Relationships
Searching the internet for images that “look Voynich” is not enough.
The most valuable comparator corpus would include manuscripts whose relationships are already known.
Source → copy.
Earlier witness → later witness.
Latin exemplar → vernacular adaptation.
Textually related manuscripts with different illustration programmes.
Visually related manuscripts with altered text.
Once those relationships are known, we can measure how real manuscript transmission changes images, labels, order, abbreviations and page design.
Those transformation distributions become much stronger controls for Voynich than isolated lookalikes.
We Need a Public Negative-Result Memory
This may be the least glamorous and most valuable infrastructure the field could build.
When a hypothesis is tested and fails, record the public conclusion.
Not every private notebook.
Not every proprietary method.
Not every half-formed idea.
But enough to tell the next researcher:
- what hypothesis family was examined;
- what public observations motivated it;
- what prediction failed or failed to generalise;
- which narrower component survived;
- what genuinely new evidence would justify reopening it.
This Part VIII ledger is an example of the principle.
A field becomes cumulative when failure becomes infrastructure.
The Ten Highest-Value Research Directions
1. Representation before language identification
Build better glyph and segmentation models. Quantify uncertainty. Determine which statistical anomalies survive different transcription assumptions.
If our basic units are wrong, every language-ranking system downstream is partly answering the wrong question.
2. A spatial document model
Encode page geometry, labels, objects, paths, foldouts and orientation. Let the model see the manuscript as a page, not merely as EVA text.
This may reveal structure invisible to linear corpus methods.
3. Cross-modal recurrence
Systematically search for repeated image motifs and test whether their textual environments recur more than chance predicts.
Do this without assigning object names first.
Internal recurrence is the cleanest route toward semantic constraints that does not require guessing an external language or species.
4. Production chronology
Integrate quire structure, hand variation, Currier regimes, ink/paint sequencing, page layout and vocabulary drift.
Do the changes form discrete production episodes?
Or a continuous evolving process?
This question may be more tractable than plaintext and could radically narrow mechanism hypotheses.
5. Material provenance
Use every non-destructive or responsibly sampled technique capable of narrowing parchment production, pigments, ink composition, protein source, microbiome, environmental history or manufacturing practice.
No single material signature is likely to say “Padua”.
Several independent material signals may narrow the geographic field.
6. Historically executable transformation models
Build ciphers, abbreviation systems, mnemonic systems and generators that a fifteenth-century scribe could actually use.
Do not optimise only for one statistic.
Require simultaneous reproduction of token structure, line effects, local vocabulary, Currier-like variation and scribal practicality.
7. Known manuscript-lineage controls
Build datasets from known copying relationships such as herbal traditions. Measure which visual features survive copying and which drift.
Then use those empirically measured transformation patterns when evaluating supposed Voynich source relationships.
8. Controlled candidate-language testing
Candidate languages should be evaluated under fixed transformation budgets.
How many insertions?
How many deletions?
How many homophones?
How many boundary shifts?
How many exceptions?
Then compare languages under the same budget rather than allowing the favourite language unlimited flexibility.
9. Semantics from repeated internal relations
Instead of guessing that a token means “root”, ask whether the token component repeatedly accompanies roots and not leaves, stars or vessels.
Then test the association in held-back cases.
A semantic class may emerge before an English gloss does.
That would be progress.
10. Archive discovery
Do not let computational glamour make us forget the possibility that the decisive clue is sitting in an archive.
A letter.
An inventory.
A sale record.
A workshop account.
A related manuscript.
A marginal note in another codex.
A provenance bridge can radically alter how the statistical evidence is interpreted. The manuscript lived in human institutions before it lived in datasets.
Research Priority Zero: Make Every Claim Addressable
Before any of those ten directions, one infrastructural habit comes first.
Every claim should be traceable to the evidence that supports it.
“The manuscript has five scribes” should point to the hand-classification study and its later critique.
“The parchment dates to 1404–1438” should point to the material dating, not to a blog repeating another blog.
“This plant is species X” should identify which morphological features support the identification and which alternatives were compared.
“This token is a keyword” should specify the corpus, transcription and metric.
“This cipher reproduces Voynich” should specify which properties, which plaintext, which parameters and which properties remain unmatched.
A field becomes more efficient when claims can be inspected without reconstructing the entire argument from scratch.
Evidence should have an address.
The Reader’s Rule: Ask What the Claim Had to Spend
Every theory spends assumptions.
Perhaps one unusual glyph becomes a null.
Perhaps one missing vowel is inserted.
Perhaps one plant is identified.
Perhaps one page is assumed displaced.
Any one of those may be reasonable.
But assumptions have a budget.
If a theory needs ten arbitrary glyph values, six inserted vowels, three rearranged word orders, four special abbreviations and two lost pages before producing its first beautiful sentence, the beauty of the sentence is not free.
The correct comparison is not:
Does this reading sound plausible?
It is:
How much freedom did the theory spend to make the reading plausible?
Good solutions become cheaper as they scale because the same rules keep paying for new text.
Bad solutions become more expensive because every new page needs another exception.
The Research Contract in One Sentence
Whatever Voynich is, the explanation must reproduce the manuscript we actually have—not a simplified version created by ignoring the features the explanation cannot handle.
That is the standard.
The rest of Part X gives readers a practical way to enforce it.
Before You Believe the Next “Voynich Solved” Headline — Ask These 20 Questions
You will see another one.
Perhaps next month. Perhaps next year. Perhaps from a university press office. Perhaps from an independent researcher. Perhaps from a machine-learning team. Perhaps from somebody who has spent twenty years on the manuscript and deserves to be heard carefully.
The headline will say some version of:
Voynich Manuscript finally decoded.
Do not reject it because we have heard that sentence before.
Do not believe it because we have heard that sentence before either.
Ask questions.
These twenty questions are designed so that an intelligent non-specialist can distinguish an interesting proposal from a manuscript-scale decipherment without needing to become a cryptographer first.
1. What exactly has been solved?
One label?
One page?
One glyph?
A transcription problem?
A language family?
A cipher mechanism?
The entire manuscript?
“Solved” is too broad to evaluate until the claim has an object.
2. Are the reading rules written down exactly?
You should be able to see how glyphs are segmented, how they transform, how spaces are handled, how uncertain forms are treated, and what the proposed phonetic, lexical or cipher values are.
If the essential rule is “I can see what the scribe intended”, the method is not yet transferable.
3. Were the rules fixed before the showcase examples were read?
A rule invented after seeing each target can explain almost anything.
A useful rule should be frozen and then applied where the researcher does not already know what they hope the page says.
4. How much of the manuscript does the method actually cover?
Ten tokens out of tens of thousands is not manuscript coverage.
A method may still be important at small scale, but its claimed scope should match its tested scope.
5. Has it been tested on material not used to build the theory?
This is the difference between explanation and fitting.
Give the method pages it did not train on.
Give it ordinary pages.
Give it damaged pages.
Give it labels.
Give it prose.
Give it a different Currier regime.
Then see whether the same machinery survives.
6. Do repeated Voynich forms receive stable readings?
If the same token appears ten times, does it preserve a recognisable lexical, grammatical or functional relationship?
Real language can be polysemous. Abbreviations can vary by context. Cipher homophones complicate direct identity.
But repeated forms cannot mean anything at all whenever the picture changes.
7. Does grammar remain grammar?
If the proposal claims a natural language, look for stable syntax, morphology, particles, agreement, function words or whatever that language should exhibit.
A dictionary without grammar is not enough for long running text.
8. What explains Currier A and Currier B?
If a proposed decoding reads both regimes identically, why do the surface distributions differ?
If it reads them differently, what rule determines the difference?
Ignoring Currier is no longer acceptable.
9. What explains handwriting variation?
Does the system work regardless of the proposed hand?
Are some apparent character changes actually scribal variants?
Does the model depend on the five-hand classification being correct?
Would it survive a continuous-hand model?
A robust decipherment should not collapse because palaeographers redraw one hand boundary.
10. What explains line and paragraph effects?
Why do certain forms prefer line starts?
Why do others prefer endings?
Why can paragraph openings be special?
If the proposed plaintext has no reason to care where the physical line wraps, what part of the encoding or writing practice creates that effect?
11. Does it explain labels as well as prose?
Labels are one of the manuscript’s best stress tests because their document function differs from long paragraphs while the script overlaps.
A method that can only read labels may have discovered a label system.
A method that can only read prose may not explain what the images are being labelled with.
12. Does the method reproduce word families, locality and predictability?
A translation can sound excellent while its production mechanism fails the statistical manuscript.
Why are near-neighbour tokens so common?
Why does vocabulary localise?
Why is character succession so constrained?
A complete account needs an answer.
13. Did the pictures generate the translation, or independently confirm it?
This question catches an enormous amount of circularity.
If a page shows a plant and the proposed translation says “this green medicinal plant…”, ask whether the same decoding rules produced that language before the researcher used the picture to choose among alternatives.
The image should ideally become an independent validator.
Not an answer key hidden inside the question.
14. Does the historical mechanism fit the fifteenth century?
Could a real person make this text by hand?
Could they maintain the rule across hundreds of pages?
Could collaborators learn it?
Does it require modern statistics, a computer-sized lookup table or knowledge unavailable in the manuscript’s historical horizon?
A modern analytical model can be valuable even if historically impossible.
But it should not be called the production mechanism without a plausible historical implementation.
15. Does the theory respect the physical codex?
Are missing leaves acknowledged?
Is uncertain page order acknowledged?
Are foldout relationships preserved?
Does the theory depend on present adjacency where codicology makes that unsafe?
The book is allowed to veto a theory about the book.
16. How many exceptions does the solution need?
Count them.
Do not let “scribal error” become invisible.
Do not let “abbreviation” become invisible.
Do not let “null” become invisible.
Do not let “different dialect” become invisible.
Every rescue clause spends flexibility.
A real mechanism may genuinely need exceptions. The question is whether exceptions are rare consequences of a stable system or the system itself.
17. Can somebody else reproduce the reading?
Give another researcher the transcription, rules and target page.
Do they obtain substantially the same result without private coaching?
If not, the method may depend more heavily on the original researcher’s intuition than its formal description admits.
18. What would falsify the claim?
A serious theory should be able to lose.
What observation would force the author to say, “this mechanism is wrong”?
If the answer is “nothing, because any mismatch can be an exception, error, null, homophone, lost page or special case”, the theory has left the evidentiary world.
19. Was it compared against serious alternatives?
A model that fits Voynich is interesting.
A model that fits Voynich materially better than plausible alternative models is explanatory.
Could a simpler generator produce the same statistical result?
Could another candidate language fit equally well with the same number of transformations?
Could a cipher family explain the visible constraint with fewer special rules?
Voynich requires comparison, not merely fit.
20. Does the result become more constrained as evidence accumulates?
This is the final question because it contains almost all the others.
A good explanation gains obligations.
The first decoded page fixes rules that constrain the second.
The second constrains the third.
Eventually the theory has very little freedom left.
A bad explanation behaves in the opposite direction. Every new page creates another special case, widening rather than narrowing the rule set.
The strongest solution is the one that becomes least able to improvise.
Five Different Things People Call a “Solution”
To reduce confusion, it helps to name the level of achievement accurately.
Level 1 — Interesting observation
Example: a previously unnoticed glyph preference, plant resemblance, layout pattern, repeated label family or archival clue.
This can be valuable research.
It is not decipherment.
Level 2 — Local reading
A small passage or label can be read plausibly under explicit rules.
This is stronger because meaning has entered.
It still needs scale.
Level 3 — Partial decipherment
A stable mechanism decodes a defined class of material—for example, a numeral system, zodiac labels, plant labels or a repeated formula—and predicts new examples within that class.
This would be a major breakthrough even if the rest remains unread.
Level 4 — Mechanism demonstration
A proposed process can reproduce important Voynich properties from known input or generate text with strikingly similar structure.
This establishes feasibility or explanatory power.
It is not the same as proving that the historical manuscript used that process.
Level 5 — Manuscript-scale decipherment
A stable, historically plausible, reproducible system explains a large majority of ordinary running text and specialised text, handles the manuscript’s major structural irregularities, generates coherent language or another demonstrated semantic representation, makes successful predictions and outperforms serious alternatives.
This is the level that deserves “Voynich solved”.
We are not there.
What Would Genuinely Change Our Mind?
A good research position is not stubborn.
It should be easy to describe what new evidence would force it to change.
Any of the following could substantially alter the current Voynich landscape.
A genuine bilingual or gloss
A manuscript witness in which a Voynich-like string is explicitly paired with readable text under a demonstrable historical relationship would transform the field.
One secure bilingual anchor can do more than thousands of speculative word matches.
A related manuscript with transparent writing
Imagine discovering another codex with unmistakably related plants, diagrams or labels, but written partly in an identified script or accompanied by explanatory text.
That could provide the missing bridge between visual tradition and textual mechanism.
A new early archival document
An inventory, commission, sale, library catalogue, workshop account or correspondence item securely datable to the fifteenth or sixteenth century could narrow provenance dramatically.
One boring archival line can be more decisive than a spectacular diagram resemblance.
A reproducible material localisation
If several independent material signals converged on a narrower production region—and equivalent comparison samples showed the signals were genuinely diagnostic—that would materially update geographic hypotheses.
A predictive image–text mapping
A token component or structural pattern that repeatedly predicts a visual category across held-back folios would be an important semantic foothold even before an English gloss is known.
A fixed decoder that works on unseen ordinary text
This is the obvious one.
Publish the rules. Freeze them. Give the decoder pages it did not use. Obtain coherent output with stable grammar and repeated meanings.
Do that across the manuscript and the field changes overnight.
Independent replication of contested palaeography
A robust independent analysis that strongly supports either discrete hands or continuous scribal drift would clarify production structure and help disentangle hand from Currier regime.
A known source or exemplar chain
If a set of images, errors, ordering relations and texts could be traced through a demonstrable manuscript lineage into Voynich, provenance and function would gain an entirely new evidentiary basis.
What Probably Would Not Change Our Mind by Itself
This list is equally important.
- One more isolated plant identification.
- One more city resemblance on the Rosettes foldout.
- One more target-language word list assembled from a handful of Voynich tokens.
- One more fluent AI-generated translation without stable decoding rules.
- One more entropy or Zipf curve considered in isolation.
- An opaque neural-network confidence score with no interpretable mapping.
- Three or four attractive translated labels.
- A theory that places missing proof on missing pages.
- A visually similar manuscript without a genealogical bridge.
- A synthetic generator that reproduces only the statistics it was tuned to reproduce.
- A high percentage “accuracy” without a defined ground truth or benchmark.
- An old attribution repeated in a new article without new evidence.
- A map, plant or jar match selected after searching one favoured region only.
- A decipherment whose rules expand every time a new folio is tested.
Any of these could be the beginning of something important.
None should be allowed to skip the middle.
How to Publish a Voynich Solution So Serious People Can Test It
If you believe you have made a breakthrough, make it easy for sceptics to try to destroy it.
That is not hostility.
That is how you give a correct idea the chance to become durable.
- State the claim narrowly. Say exactly what you believe you have decoded, identified or explained.
- Name the source data. Which transcription? Which manuscript images? Which folios? Which edition or dataset?
- Publish the representation decisions. Explain glyph segmentation, uncertain readings, spaces and special forms.
- Publish the rules before the showcase. A reader should know the mechanism before seeing the examples chosen to flatter it.
- Show ordinary cases. Do not publish only the spectacular three.
- Show failures. If twenty eligible cases exist and the rule works on eleven, say eleven of twenty.
- Keep the denominator. A success rate without the number of opportunities for failure is nearly meaningless.
- Use held-back material. Let the final rules face pages they did not help explain.
- Compare alternatives under equal freedom. If Italian receives ten transformations, give Latin, French, German or other candidates the same budget where appropriate.
- Count exceptions. Nulls, scribal errors, special abbreviations and boundary shifts should not disappear into prose.
- Make code or procedures available where feasible. Another researcher should be able to reproduce the computational result.
- Separate observation from interpretation. “This token clusters here” and “this token means root” are two claims.
- Separate possibility demonstrations from identity claims. A mechanism that can produce Voynich-like text is not yet the mechanism Voynich used.
- State uncertainty. Damaged glyphs, ambiguous labels and doubtful plant identifications should remain visibly uncertain.
- State falsifiers. Tell readers what result would make you abandon or revise the theory.
- Use the right word for the achievement. Observation, local reading, partial decipherment, mechanism demonstration and manuscript-scale decipherment are not synonyms.
If a proposed solution survives that treatment, it will emerge stronger.
If it does not survive, the failure will still leave the field with usable knowledge.
Publish so that the idea can lose. If it refuses to lose, we may finally have something.
Frequently Asked Questions
Has the Voynich Manuscript been solved?
No manuscript-scale decipherment has achieved broad scholarly acceptance under stable, reproducible rules that explain the major textual and structural constraints discussed in this article.
Do we know what language it is?
No. Natural-language hypotheses remain serious, but no underlying language has been established to consensus.
Does the text contain meaning?
That remains unresolved. The visible text is highly structured, non-uniform and rich in long- and short-range regularities. Those properties are compatible with meaningful language or encoded meaning, but structured non-semantic or partly semantic generation has not been eliminated simply by the existence of structure.
Could it be a cipher?
Yes. Simple substitution is inadequate as a blanket model, but richer historically plausible cipher families remain viable. The 2025 Naibbe demonstration is important precisely because it shows that meaningful Latin or Italian can be transformed into ciphertext sharing several Voynich-like statistical properties without proving that Naibbe itself was the historical system.
Could it be generated nonsense?
Some rule-generated non-semantic models remain possible. But generated-looking surface structure does not establish meaninglessness, because meaningful text can also be transformed into highly constrained output.
Is it a hoax?
“Hoax” needs decomposition. A simple twentieth-century forgery by Wilfrid Voynich is contradicted by the medieval material evidence and seventeenth-century documentary history. A genuinely medieval text that is partly or wholly non-semantic is a different hypothesis and remains logically open.
Was it made in Italy?
Northern Italian manuscript culture is a historically plausible comparison environment, and several relevant comparator traditions exist there. That does not establish Italy, Padua or another specific location as the manuscript’s proven place of production.
Were there five scribes?
A five-hand model published by Lisa Fagin Davis in 2020 remains influential. A 2026 critique by Torsten Timm argues that some of the variation may be continuous and compatible with fewer distinct hands. The secure conclusion is that meaningful handwriting variation exists; the exact number of scribes is contested.
Do we know what the plants are?
Many identifications have been proposed, and some individual images may eventually be identifiable with useful confidence. There is no complete secure plant inventory, and single-feature resemblance is usually too weak to carry a species identification alone.
Is the Rosettes foldout a map?
It may be map-like, cosmological, processual or another kind of relational diagram. What is secure is that it deliberately organises nine major regions within a connected spatial system. No specific city or geography is established.
What are the women in the pools?
The figures, liquids, containers and tubes of Quire 13 form a repeated networked visual system. Balneological, anatomical, reproductive, medical, cosmological and process-based interpretations have all been discussed. The exact function remains unresolved.
Are the starred pages recipes?
“Recipe” is a useful conventional description because the pages contain many short segmented entries, often associated with star-like marginal markers. The secure observation is the record-like segmentation; procedural recipe semantics remain unproved.
Can AI solve Voynich?
AI can accelerate corpus search, handwriting comparison, image retrieval, anomaly detection, simulation and stress-testing. It cannot replace explicit rules, historical plausibility, independent prediction and replication. A fluent AI translation is not evidence by itself.
What is the most important thing we know about the text?
That it has structure at many interacting scales: within glyph combinations, within visible tokens, at line and paragraph boundaries, across local vocabulary, across Currier regimes, across document functions and across the manuscript’s physical and visual geography.
What would a real solution look like?
Stable. Explicit. Historically plausible. Predictive. Reproducible. Able to explain ordinary pages, not just spectacular ones. Able to survive Currier variation, line effects, labels, word families, image relationships and codicology. And increasingly boring as the same rules keep working.
Where should a new reader begin?
Begin with this master. Then follow the specialist articles in the reading map below according to the question that interests you. Do not begin by choosing a language or a solution. Begin by learning which constraints that solution will have to survive.
If You Read Nothing Else, Keep These Twenty Lines
- The Voynich Manuscript is a genuine medieval parchment codex.
- Its sampled parchment belongs to an early-fifteenth-century radiocarbon horizon.
- The surviving book is incomplete.
- Present page order is evidence, not guaranteed original process order.
- The primary script remains undeciphered.
- EVA is a transcription convention, not a phonetic solution.
- The visible writing is highly constrained rather than randomly distributed.
- Line and paragraph position affect the text.
- Currier A and B are robust textual regimes whose cause remains unknown.
- Handwriting variation is real; the exact number of scribes is contested.
- Voynich tokens form dense families of near-related forms.
- Vocabulary localises across the manuscript.
- Labels are an important functional text class.
- The manuscript contains several coherent visual regimes.
- Text and image interact materially on the page.
- Historical comparators show that plants, medicine, celestial timing, baths, drugs and recipes could coexist in medieval knowledge environments.
- Natural language remains possible.
- Ciphering or another transformed meaningful representation remains possible.
- Generated or partly non-semantic structure also remains possible.
- No accepted manuscript-scale solution currently earns the right to ignore the surviving constraints above.
That is where a new reader should start.
Not from zero.
Complete Voynich Research Library
This master is the canonical doorway. The specialist articles below are grouped by research job rather than publication date. The library is intentionally open-ended: a new node belongs here only when it earns a distinct evidentiary job rather than repeating an existing owner.
Foundations: Evidence, Classification and Provenance
- What We Actually Know
- The Manuscript Before the Mystery
- What the Writing Does Before We Know What It Says
- What the Pictures Can and Cannot Tell Us
- The Broken Provenance Chain
- The Six Sections That May Not Be Six Sections
- Ink, Pigment, Parchment and the Physical Evidence
Physical Production, Source History and Human Hands
- The Missing Leaves and Quire Reconstruction
- Quire 8: The Transition Zone Where the Manuscript Stops Lining Up
- Marginalia and the Later Hands
- f116v: The Last Page That Almost Looks Readable
- The Scribe Problem
- The Order of Making: Drawing, Writing, Paint and Assembly
- The Bifolium as a Production Unit
- The Foldouts: When One Page Was Not Enough
- When Text Meets Image: How the Page Was Planned
- The Unruled Page: Ruling, Pricking, Wavy Lines and Scribal Control
- The Geometry Tools Problem: Compass, Straightedge and How the Diagrams Were Constructed
- The Almost-Correctionless Manuscript: Errors, Emendations and Scribal Control
- The Exemplar Problem: Was Voynich Copied From Something Else?
The Visual World
- The Herbal Pages
- The Zodiac Pages
- The Astronomical and Cosmological Pages Beyond the Zodiac
- The Rosettes Foldout
- The Human Figures, Pools and Tubes
- The Vessels and Plant Fragments
- The Starred Pages
Transcription, Characters, Direction and Boundaries
- EVA, Transcription and the Segmentation Problem
- The Alphabet Problem
- Abbreviation or Alphabet?
- Rare Glyphs
- The Minim Strings
- The Gallows Characters
- The Bench Characters
- The Ductus Problem: Pen Lifts, Stroke Order and Where a Glyph Really Begins
- Spaces and Word Boundaries
- Which Way Does Voynich Read? Direction, Rotation and the Geometry of Sequence
- Where Does a Sentence End? Punctuation, Paragraphs and the Missing Boundary Problem
- The Key-Like Sequences: f49v, f57v, f66r and the Characters That Refuse to Behave Like Prose
- f57v: The Seventeen-Sign Ring That Looks Like a Key
How Voynichese Behaves
- Currier A and B
- The Line as a Unit
- The Grove Words: Why Paragraph-Initial Words Grow Gallows
- Word Families
- Why Voynichese Is So Predictable
- Frequency Laws: Zipf, Vocabulary Growth and the “Looks Like Language” Trap
- Labels Versus Running Text
- Keywords Without Meanings
- The Q-Series
- Exact Repetition
- Syntax Before Semantics
- Where Are the Numbers? Numerals, Quantities, Counts and Coordinates
Mechanisms and Decipherment Standards
- Cipher, Plaintext or Generated System?
- The Crib Problem: Why Voynich Has No Secure Rosetta Stone
- What a Real Voynich Decipherment Must Survive
Research Method, Controls and History
- The Control Problem: What Should Voynich Be Compared With?
- A Century of Voynich Research: How the Question Changed From Newbold to AI
Use this library as a map of distinct questions, not as a list of mutually exclusive answers. Structure is evidence; structure is not translation.
Canonical Research Extensions — Navigation, Paratext, Representation and Drift
These specialist owners extend the master into narrower evidence layers that become visible only after the broad manuscript architecture is already protected. They remain separate because each asks a different question and each has a different falsification burden.
Navigation, Later Layers and Reconstructed Order
- The Month Names Problem — readable Romance zodiac annotations as a later reception layer rather than a direct key to Voynichese.
- The Two Numbering Systems — quire marks and foliation as separate historical snapshots of the codex.
- Quire 9: The Foldout That Remembers an Earlier Binding — stitching, fold geometry and numbering that preserve evidence of rebinding.
- The Pharmaceutical Order Problem — three fragment-and-container bifolia whose visual sequence may preserve an older order than the present binding.
- The Paratext Problem — the missing readable title, author statement, headings, contents and original interface of the book.
Representation, Phonology and Continuous Change
- The Vowel Problem — whether the visible system exposes, suppresses or transforms vowel information and why no simple vowel class has been recovered.
- The Word-Length Problem — the narrow token-length envelope and how segmentation, spaces, abbreviation, language, cipher and generation can all affect it.
- The Drift Problem — whether textual change is better understood as discrete A/B boxes, continuous movement, or several hidden variables acting together.
Later navigation is not original semantics. Statistical continuity is not chronology. A measurable representation effect is not yet a language identity.
Deep Research Routes Across the Voynich Estate
The master now connects not only to the original specialist suite but also to the newer high-resolution text-system work. These routes are designed for readers who want to move through the evidence in a disciplined order rather than bounce among isolated theories.
Newest text-system nodes
- The Alphabet Problem — how many Voynich characters are there before we even ask what they mean?
- The Minim Strings — repeated strokes, segmentation and the danger of counting what may be components rather than letters.
- Rare Glyphs — one-off signs, ligatures, scribal variants and the difference between anomaly and alphabet expansion.
- Abbreviation or Alphabet? — how medieval scribal compression changes what a visible character sequence might represent.
Route 1 — Object → representation → decipherment
Manuscript Before the Mystery → EVA and Segmentation → Alphabet Problem → Minim Strings → Rare Glyphs → Abbreviation or Alphabet? → Syntax Before Semantics → Real Decipherment.
Route 2 — Image → label → recurrence → historical control
Pictures Can and Cannot Tell Us → Labels Versus Running Text → Keywords Without Meanings → Carrara Herbal → Masson 116. The route moves from internal relation to external comparison without turning resemblance into provenance.
Route 3 — Distribution → model → held-out test
Currier A and B → Word Families → Exact Repetition → Predictability → Statistical Model Selection → Cross-Validation. The reader moves from an observed pattern to competing explanations and finally to testing on data the explanation did not fit.
Hidden conceptual bridges worth following
- How Scientific Research Works — why Voynich progress depends on separating exploration from confirmation, preserving negative results and allowing the world to say no.
- The Model Is Part of the Message — a strong bridge to abbreviation, verbose encodings and the cost of explanatory machinery.
- Compression and Reconstruction — a reminder that a compact representation is only useful if a decoder can reconstruct what the representation was intended to preserve.
- Preservation Masters and Access Copies — why the manuscript image, transcription and analytical dataset are different evidence layers.
- The Museum Label Is Tiny Scholarship — a useful parallel for Voynich labels: proximity and compression constrain a label’s job without automatically giving its wording a name.
- Why a Few Marks Become a Face — the visual-recognition bridge explaining why human perception can complete ambiguous plants, bodies and diagrams too confidently.
- What Is Robustness? — a model of what a real decipherment should do when the folio, section, hand or example changes.
- When Information Survives but Knowledge Dies — the deepest historical connection: a representation can survive intact after the human knowledge needed to reconstruct it has disappeared.
Deep connection is useful only when it narrows the next question. A surprising analogy that creates no test remains an analogy.