Voynich Longform 02 · KNOWLEDGE
What Is the Voynich Manuscript? | Everything We Can Responsibly Say
Before asking what the Voynich Manuscript means, it is worth asking a quieter question.
What is actually there?
That question sounds elementary. It is not. The Voynich Manuscript has accumulated more than a century of modern fame, proposed decipherments, imagined origins, linguistic identifications, botanical claims, cryptographic systems, historical associations and cultural mythology. Once a mystery becomes famous, it becomes difficult to see the object without also seeing the stories attached to it.
This article deliberately works in the opposite direction. It begins with the surviving codex, then adds only what can be supported at the appropriate level of confidence. Where evidence is strong, we will say so. Where interpretation remains open, we will keep it open. Where a familiar label is useful but provisional, we will mark it as such.
The governing question is simple:
What can we responsibly say about the Voynich Manuscript before a preferred solution begins steering the evidence?
This is the KNOWLEDGE article in the new Voynich explanatory sequence. The first article, Voynich Manuscript | Why Humans Need to Solve a Mystery, examined the human receiver: pattern recognition, premature certainty, expertise, confidence and the desire for closure. Here we remove that narrative lens as much as possible and stabilise the factual floor.
1. The shortest responsible answer
The Voynich Manuscript is a handwritten illustrated codex now held by Yale University’s Beinecke Rare Book and Manuscript Library as Beinecke MS 408. It is written in an unidentified script, its text has not been generally accepted as deciphered, its author or authors are unknown, and its illustrations include plants, circular astronomical or astrological-looking diagrams, human figures associated with pools or channels, pharmaceutical-looking containers and plant parts, and pages of mostly unillustrated text marked by star-like forms.
Scientific and codicological study constrain important parts of its history without resolving its meaning. The parchment belongs to the medieval or early Renaissance material world; ink and pigment analyses are broadly consistent with historical manuscript production rather than a simple modern fabrication. The manuscript’s later ownership history is substantially documented, especially from the seventeenth century onward, while earlier production circumstances remain uncertain.
The text is not visually random. It exhibits recurring glyphs, recurring token forms, positional tendencies, page and section differences, and statistical structure that can be measured independently of translation. Researchers have identified textual subpopulations historically labelled Currier A and Currier B, and palaeographic work has argued for multiple scribal hands. These findings constrain explanation but do not by themselves identify the language, encoding system, purpose or semantic content.
That is already a great deal of knowledge.
It is also much less than a decipherment.
2. Why the physical manuscript comes first
The strongest evidence in any manuscript problem begins with the artefact itself. Language theories can change. Historical narratives can be revised. Proposed plant identifications can come and go. The physical object places constraints on all of them.
This priority is developed throughout our earlier research, especially The Manuscript Before the Mystery. The basic rule is that a theory should adapt to the object, not the object to the theory.
The surviving codex is made from parchment and is bound as a book. Its pages contain writing and illustrations executed in inks and pigments of varying appearance. Some leaves are missing. Some foldouts create unusually large working surfaces. The order in which we see pages today reflects the surviving bound object, but manuscript scholars must still ask how original production, later binding, missing material and possible rearrangement affect interpretation.
Yale’s Beinecke Library describes the codex as approximately 240 pages long and physically about 23.5 centimetres high by 16.2 centimetres wide, with a thickness of about 5 centimetres in a 2009 materials report. Exact page counts can vary depending on whether one counts folios, sides, foldouts and missing leaves, so responsible descriptions should specify the counting convention rather than treat one number as metaphysically exact.
This physicality matters because manuscript meaning is partly constrained by manufacture. A text written continuously across a fold behaves differently from independent labels. Ink laid down before a fold was cut or repaired can reveal sequence. Quire structure can constrain page relationships. Pigment placed over writing or vice versa can affect chronology. The book is not merely a container for symbols. Its construction is evidence.
3. Parchment dating: what it tells us and what it does not
Radiocarbon dating of the parchment is among the most important constraints on the manuscript. It places the animal-skin writing support in the early fifteenth century, commonly reported as a range around 1404–1438 for sampled leaves.
This is strong evidence about the parchment.
It is not identical to a precise date for every act involved in producing the manuscript.
Radiocarbon dating estimates when the animal from which the parchment was made died. Parchment could in principle be stored before use, although storage over very long periods would require additional explanation. Writing, colouring, binding, rebinding and later annotations can occur at different times. The dating therefore creates a powerful chronological constraint without magically revealing the exact year in which a scribe wrote a particular line.
This distinction is essential because popular accounts often compress “parchment dates to the early fifteenth century” into “the manuscript was written in year X.” The former is materially grounded. The latter requires additional inference.
Good historical reasoning stacks constraints. If parchment dating, ink analysis, palaeography, artistic conventions and documented later provenance all point toward compatible historical windows, confidence increases. If one line of evidence conflicts strongly with another, the contradiction needs explanation.
The value of the radiocarbon result is therefore not that it solves authorship or language. It prevents large classes of chronology from being accepted casually. Any proposed production story must fit the physical age of the support or provide a plausible reason why older parchment was used later.
4. Ink and pigment: material consistency, not translation
The Voynich Manuscript has also been examined using materials analysis. A McCrone Associates report prepared for the Beinecke in 2009 analysed inks and pigments from the manuscript. Such work can identify classes of material and compare them with substances used historically.
The strongest responsible conclusion is modest: the materials do not present the simple signature of a modern object fabricated with obviously anachronistic inks and pigments. They are broadly compatible with historical manuscript materials.
That does not tell us what the text means. It does not establish a particular city. It does not identify a language. It does not prove that every colour was applied in the same production phase. It does not establish a named author.
Materials evidence has a particular job: constrain what could physically have happened.
This is a recurring principle across the Voynich project. Each discipline should be allowed to answer the questions it can actually own. Chemistry can characterise material. Codicology can analyse construction. Palaeography can examine writing hands. Statistics can measure distributions. None of those fields should be forced to produce semantics when semantics is outside the evidence they directly control.
This separation protects strong findings from being weakened by speculative extensions. “The ink composition is compatible with historical manuscript practice” can remain useful even if a proposed geographic theory built on top of it later fails.
5. The writing system: what we can observe
The text is written in a script that has not been securely identified with an ordinary known alphabet. Researchers therefore often work with transcription systems that assign conventional Latin characters or symbols to recurring Voynich glyph forms. Those transcriptions are analytical tools, not translations.
This distinction matters enormously.
If a transcription writes a Voynich glyph as “q”, that does not mean the glyph is the Latin letter q, represents the sound /k/, or has any particular linguistic value. The symbol is a label used so that researchers can compare sequences computationally and communicate about recurring forms.
The script contains a relatively limited set of frequent glyph shapes alongside rarer forms and ambiguous compounds. Some glyphs are visually simple. Others have tall or elaborate forms traditionally nicknamed “gallows” because of their appearance. Spacing creates units that researchers often call words or tokens, but the semantic status of those spaces is not guaranteed. A visual word boundary may correspond to a linguistic word, an encoded unit, a production convention or something else.
What can be said more safely is that glyph sequences show constraints. Some forms prefer particular positions. Some combinations recur frequently. Some possible-looking combinations are rare or absent. Certain token families differ by small systematic changes. These patterns are measurable and form the basis of much of the statistical research in the archive.
They are among the reasons simplistic claims that the text is merely arbitrary scribbling face a high evidential burden. But structured output can arise from multiple kinds of mechanisms, including natural language, encipherment, abbreviation systems, formulaic notation, constrained generation and mixtures. Structure narrows the field without selecting one answer automatically.
6. Transcription is a model of the page
Because the original text is unfamiliar, every machine-readable transcription makes decisions.
Are two visually similar marks the same glyph or different glyphs? Is a ligature one unit or two? Does a damaged mark belong to one class? Should uncertain readings be represented as alternatives? How should line breaks, paragraph breaks, labels and marginal elements be encoded?
These are not trivial technical details. Statistical results depend on representation.
If a transcription system splits one complex sign into two characters, entropy and sequence statistics may change. If ambiguous glyphs are normalised differently, token families may change. If uncertain spaces are resolved inconsistently, word-length distributions may shift.
This is why strong quantitative claims should state which transcription, preprocessing and unit definitions were used. Results that remain stable across reasonable transcription choices deserve more confidence than results that appear only under one fragile representation.
The broader principle is familiar from data science: raw data is rarely truly raw. Measurement, classification and encoding decisions shape the dataset before analysis begins.
For eduKateAI, this becomes a provenance rule. A numerical result should not be stored merely as “Voynich token entropy = X.” It should retain the corpus, transcription version, tokenisation rule, exclusions and comparator definition that produced X. Without those dependencies, later retrieval can turn a conditional measurement into a false universal fact.
7. The text has internal structure
One of the safest broad statements about Voynichese is that its surface text is internally structured.
Glyph frequencies are uneven. Token frequencies are uneven. Some glyphs strongly prefer certain positions. Recurrent token families exist. Local context matters. Line beginnings and endings can exhibit different tendencies. Some pages and sections have measurably different vocabularies or distributions.
This is why our earlier research could investigate Word Families, exact repetition, entropy, positional effects and syntax-before-semantics without claiming a translation.
Internal structure is evidence because proposed mechanisms must reproduce it.
A simple random generator with independent characters is unlikely to produce the same combination of local constraints, token families and page-level variation unless those properties are deliberately built into the generator. A proposed cipher operating on ordinary language must explain why the resulting surface statistics have the observed form. A proposed abbreviation system must account for recurring transformations. A natural-language hypothesis must explain unusual features rather than merely pointing to language-like ones.
The responsible claim is therefore neither “the statistics prove it is language” nor “the statistics prove it is meaningless.”
The responsible claim is: the observed surface system is constrained enough that serious hypotheses should be tested against its measurable structure.
8. Word families and near-neighbour forms
Readers quickly notice that many Voynich tokens resemble one another. One form may differ from another by a single initial glyph. Another may contain an added final element. Similar strings can appear in clusters.
Researchers often call these word families because the visual relation resembles morphological families in known languages. That analogy is useful descriptively but should not be promoted to semantics prematurely.
Several mechanisms can produce near-neighbour token families.
Natural language morphology can add prefixes, suffixes and inflections. Abbreviation systems can produce families around repeated stems or formulae. Copying with local modification can generate similar forms. Cipher mechanisms may preserve or transform structural relationships. Constrained generative systems can deliberately build tokens from reusable components.
The presence of families therefore constrains but does not identify mechanism.
What makes the phenomenon particularly interesting is its interaction with local context. If similar words tend to occur near one another, on particular pages or in related positional environments, the mechanism must account for that clustering. A translation theory that treats each token independently may miss important dependencies.
This is one place where the next MECHANISM article will go further. The KNOWLEDGE article only needs to stabilise the observation: near-neighbour token families are a real feature of the surface corpus, while their semantic interpretation remains open.
9. Exact repetition and formulaic behaviour
Exact repetition also occurs, but its distribution matters more than the mere fact that it exists.
Human language repeats common words. Technical manuals repeat phrases. Recipes repeat instructions. Liturgical texts repeat formulae. Tables repeat labels. Generated systems repeat templates. Copying processes repeat source material.
So “the manuscript repeats itself” is not a diagnosis.
The research question is how much repetition occurs, at what scale, in which contexts, and how the observed profile compares with appropriate controls.
This is a good example of why the Control Problem is foundational. A feature cannot be called unusual without a baseline. Comparing Voynich only with modern literary prose may produce one conclusion; comparing it with medieval recipe collections, technical tables, abbreviation-heavy manuscripts or cipher texts may produce another.
Matched controls are difficult because no comparison object shares every property at once. Period, language, genre, script, scribal practice, length, layout and subject all influence text statistics. The absence of a perfect control does not make comparison impossible, but it should make claims proportional.
Knowledge here therefore consists partly of method: repetition is measurable; interpretation requires carefully chosen comparison.
10. Entropy and predictability
Voynichese has often attracted attention because some measures suggest unusual predictability relative to certain natural-language comparators. In simple terms, knowing preceding glyphs can constrain what is likely to come next more strongly than one might expect under some comparison sets.
This is scientifically interesting because predictability is a clue about mechanism.
But entropy values are not magic fingerprints. They depend on alphabet definition, transcription, normalisation, sample length, conditioning order and comparator corpus. A low or high value cannot be interpreted responsibly without those choices.
High predictability may arise from strong orthographic constraints, repetitive formulae, restricted vocabulary, abbreviation, generative templates, cipher transformations or other mechanisms. Natural languages themselves vary substantially across scripts, genres and preprocessing choices.
The durable conclusion is that Voynich text offers measurable information-theoretic structure worth explaining. The mechanism article will ask how candidate systems might generate it. The adversarial article will ask which candidate explanations fail when forced to reproduce several statistics simultaneously.
The KNOWLEDGE layer does not choose the winner. It preserves the measurement problem correctly.
11. Currier A and Currier B
A major observation in Voynich research is that the text is not homogeneous across the entire manuscript. Prescott Currier identified patterns that led to the conventional labels Currier A and Currier B for two broad textual populations.
The labels are useful because they point to reproducible differences in token usage and other textual tendencies.
What they mean remains unresolved.
They could reflect different scribes, different stages of production, different content, different linguistic varieties, different encoding conventions, different source materials or combinations of factors. Later palaeographic work identifying multiple hands adds another layer without collapsing the possibilities into one explanation.
It is therefore useful to preserve two propositions separately:
- There are measurable textual populations conventionally called Currier A and B.
- The causal explanation for those populations is not securely established.
This separation is a model of how the entire KNOWLEDGE article should operate. Stable observation first. Competing explanations later.
12. Multiple scribal hands
Palaeographic analysis has argued that several distinct scribal hands contributed to the manuscript. This matters because the Voynich Manuscript should not automatically be imagined as one person sitting alone and writing every page in one continuous act.
Multiple hands can imply collaboration, sequential work, copying, workshop production or other social processes. But again, hand identification does not automatically identify authorship, language or purpose.
A scribe is the person who writes. The intellectual author of content may or may not be the same person. One author can be copied by multiple scribes. Several contributors can work under a shared plan. A compiler can combine sources. Later additions can occur.
The distinction is especially important for computational analysis. If different hands correlate with different token distributions, some variation may arise from scribal practice. If Currier groups cross hand boundaries, other explanations gain weight. These relationships can be tested without knowing semantics.
For eduKateAI, this becomes an ontology lesson: PERSON, SCRIBE, AUTHOR, COMPILER and OWNER are not synonyms. A robust knowledge graph should preserve role distinctions even when the same historical person could occupy several roles.
13. The illustrated sections
The Voynich Manuscript is commonly divided into sections based largely on illustration type and layout. Yale’s public description groups them into six broad categories: plant specimens; astronomical and astrological drawings; balneological drawings involving baths or bathing; cosmological medallions; pharmaceutical drawings of herbs, roots and containers; and mostly unillustrated text often described as recipes because star-like markers introduce short entries.
These section names are useful navigational shorthand.
They are not translations.
Calling a section “herbal” means the pages visually resemble the conventions of plant-centred manuscripts. Calling another “pharmaceutical” reflects the presence of plant parts and vessel-like forms. Calling a group “balneological” reflects figures appearing in liquids, pools or connected structures. The labels organise visual evidence while carrying varying degrees of genre inference.
The safe method is to retain the conventional names for communication while remembering that the original users may have categorised the material differently.
Section boundaries matter because text statistics can change with visual section. That suggests the writing and imagery are not independent decoration. Yet correlation alone does not tell us whether text differences reflect subject vocabulary, production stage, scribe, encoding regime or another factor.
14. The herbal pages
The herbal pages are among the manuscript’s most immediately recognisable and most treacherous features.
Large plant-like illustrations occupy substantial portions of the pages, accompanied by text. Many drawings contain recognisable botanical components: root-like masses, stem-like structures, leaf-like forms and flower-like elements. Some appear plausible as stylised plants. Others combine features in ways that make exact species identification uncertain.
Researchers have proposed many plant matches. The difficulty is distinguishing genuine identification from resemblance.
Medieval botanical illustration was not always naturalistic. Copying across manuscript generations could transform images. Artists might emphasise diagnostic, symbolic or useful parts rather than visual realism. Conversely, a strange drawing could reflect composite construction, damaged exemplars, misunderstood copying or something that is not intended as a literal plant portrait.
This is why What the Pictures Can and Cannot Tell Us remains foundational. Images are evidence, but they do not come with their intended labels attached.
A strong plant identification would benefit from independent convergence: multiple morphological features, historical plausibility, comparison with relevant manuscript traditions and ideally textual evidence that was not derived from the plant assumption itself.
15. Astronomical and astrological-looking pages
Another group contains circular diagrams, stars and zodiac-associated imagery. Some folios include recognisable zodiac motifs, which provide stronger broad category anchors than many of the plant identifications.
Yet recognition of a zodiac sign does not translate surrounding labels or establish the exact intellectual tradition in which the diagram was used.
Medieval and Renaissance astronomy, astrology, medicine, calendrical practice and cosmology were deeply interconnected. Similar visual forms could serve different functions across manuscripts. Zodiac imagery could relate to calendrical cycles, medical astrology, natal astrology, timing, body correspondences or other systems.
The responsible claim is therefore layered. Some imagery belongs recognisably to the wider historical visual world of zodiac and celestial representation. The exact function of each Voynich diagram and the meaning of adjacent text remain unresolved.
This is a good place to resist a common inferential leap: identifying one visual convention does not identify the entire manuscript. A technical compendium can borrow imagery from several traditions. A copied diagram can travel far from its origin. A later compiler can combine sources.
16. The balneological or biological-looking pages
Some of the most distinctive pages show numerous human figures, many apparently female, associated with pools, channels, containers or tube-like structures. These pages have generated interpretations involving bathing, medicine, anatomy, reproduction, waters, spas, alchemy and symbolic systems.
The illustrations justify descriptive claims: human figures appear in liquid-like or vessel-like environments connected by structures.
The purpose of those scenes is less certain.
Calling the section “balneological” is conventional because bathing and water imagery offers a plausible historical analogy. Calling it “biological” reflects another interpretive tradition. Neither title should be mistaken for a decoded chapter heading.
These pages are particularly valuable for comparative work because their compositions are unusual enough to invite targeted search. But unusualness increases the risk of overvaluing partial matches. A strong comparison should account for composition, chronology, iconographic function and the prevalence of similar motifs elsewhere.
Rainbolt-style clue ranking is useful here: the best parallel is not the most visually exciting resemblance but the one that most sharply reduces plausible contexts after controls are considered.
17. Cosmological diagrams and foldouts
The manuscript contains elaborate circular and multi-part diagrams, including foldouts that provide much larger working surfaces than ordinary folios. These have often been called cosmological because they resemble historical diagrammatic attempts to represent worlds, heavens, elements, geography or abstract systems.
One especially famous foldout is often informally called the Rosettes folio because of its linked circular structures.
Its visual complexity has generated geographic, architectural, cosmological and symbolic interpretations.
Knowledge discipline requires us to distinguish the layout from the reading. We can describe linked circular regions, connecting forms, textures, towers or crenellation-like details and central structures. Determining whether those are cities, cosmological spheres, maps, diagrams of processes or symbolic constructs requires further evidence.
Foldout format itself is meaningful evidence because it suggests the maker needed a surface larger than a standard page for certain diagrams. That constrains function at a very general level: the relationships displayed across the expanded sheet likely mattered enough to justify special construction.
Again, construction gives us something before interpretation does.
18. The pharmaceutical-looking section
Other pages contain smaller plant parts alongside objects that resemble jars, vessels or containers. This combination has encouraged the conventional label “pharmaceutical.”
The historical analogy is reasonable because medieval medical and pharmaceutical manuscripts often organised materia medica around plants, preparations and containers.
But resemblance to a pharmaceutical manuscript is not equivalent to proof of pharmaceutical function.
The vessels may encode categories, storage forms, decorative conventions or objects with functions we do not understand. The plant parts may be ingredients, identifiers or schematic references. The text might name, describe, instruct, classify or do something else entirely.
The section is nevertheless important because it potentially links the manuscript’s plant imagery to a practical knowledge system. If the same visual motifs, token families or scribal patterns connect herbal and pharmaceutical pages, those relationships can be analysed without knowing exact semantics.
Such cross-section relationships are precisely the kind of evidence the future ATLAS should make easy to traverse.
19. The recipe-like pages
Near the end of the manuscript are pages dominated by short textual paragraphs or entries, often introduced by star-like or flower-like markers. They are commonly called the recipe section because the layout resembles lists of short instructions or preparations found in practical manuscripts.
Again, the label is descriptive analogy rather than translation.
The layout itself is useful evidence. Repeated markers suggest segmentation into entries. Entry lengths can be measured. Repeated opening or closing forms can be compared. Vocabulary overlap with other manuscript sections can be analysed. Scribal hand can be considered.
If the section really contains recipes, we might expect certain structural properties: recurring ingredient terms, action formulae, quantities, preparation verbs or ordering patterns. But because we do not know the language or units, those expectations must be translated into testable surface predictions carefully.
The important knowledge point is that visual layout implies internal organisation. The manuscript is not a continuous undifferentiated stream. Page architecture itself supplies structure.
20. Text and image relationships
One of the great unresolved questions is how the text relates to the images.
In some places, short strings appear positioned like labels. Elsewhere, paragraphs accompany large illustrations. Diagrammatic pages place text around or within visual structures. Recipe-like entries rely much less on imagery.
These layouts make it reasonable to infer that text and image are related in at least some contexts.
But relation has many forms.
A label can name an object. A caption can describe it. A technical paragraph can explain its use. A mnemonic text can cue information not depicted directly. A decorative image can mark a section without encoding local semantics. A diagram can organise concepts while the surrounding text discusses processes rather than labels.
This is why image-first decipherments must be handled carefully. If we assume an image identity and then force nearby text to yield the expected word, the apparent convergence may be circular.
A stronger approach seeks independent relationships. Do tokens associated with visually similar objects recur systematically? Do labels exhibit distinctive structural classes? Does a proposed semantic mapping generalise to images not used to build it?
The pictures can guide questions. They cannot automatically serve as an answer key.
21. What we know about provenance
The manuscript’s documented ownership history becomes much firmer in the early modern period than at its point of production.
Yale’s current public account describes an early chain associated with the court of Holy Roman Emperor Rudolf II, the pharmacist Jacobus Horčický de Tepenec, Georg Baresch, Johannes Marcus Marci and the Jesuit scholar Athanasius Kircher. A 1666 letter from Marci to Kircher survives with the manuscript and is one of the most important provenance documents.
Some details in the early provenance tradition depend on later reports and should not be treated with the same confidence as the surviving manuscript or letter. For example, historical claims about Rudolf II purchasing the manuscript and believing it to be by Roger Bacon belong to the provenance tradition, but the Roger Bacon attribution itself is not accepted evidence of authorship.
The modern custody chain is clearer. Wilfrid Voynich acquired the manuscript in 1912 from a Jesuit collection near Rome. After his death it passed through Ethel Voynich and Anne Nill, then was sold to bookseller H. P. Kraus in 1961. Kraus later donated it to Yale’s Beinecke Library in 1969.
This chain is more than biography. Provenance constrains what the object could have encountered, how later annotations might have arisen, and which archival trails deserve investigation.
But provenance after the sixteenth or seventeenth century cannot by itself reveal where the manuscript was first made.
22. Wilfrid Voynich and the modern mystery
The manuscript takes its modern popular name from Wilfrid M. Voynich, the antiquarian bookseller who acquired it in 1912.
Voynich’s acquisition transformed the object’s modern research history because he brought it to the attention of scholars and circulated photographs and information about it. The manuscript soon attracted cryptographers, historians and linguists.
The modern mystery therefore has two chronologies.
The first chronology is the manuscript’s historical production and custody.
The second is the history of attempts to explain it.
These should not be confused. A twentieth-century theory about medieval authorship is evidence about modern interpretation, not direct evidence about medieval production. A famous attribution can become culturally durable without becoming materially stronger.
This distinction becomes especially important when studying the research literature. Claims need dates and provenance just as manuscripts do. Who proposed the idea? On what evidence? Was the claim later revised? Did later writers repeat it accurately?
A knowledge system that cannot track the history of its own claims risks mistaking old speculation for old fact.
23. The Roger Bacon attribution
Roger Bacon has long hovered around the manuscript’s story because an early provenance tradition associated the book with him. That attribution helped make the manuscript attractive to cryptographic and historical speculation.
But the parchment dating to the early fifteenth century is incompatible with the manuscript having been physically produced by Roger Bacon, who lived in the thirteenth century, unless one proposes an extraordinary chain involving much later copying of an earlier work.
Such a copy hypothesis is logically possible but requires evidence. It should not be smuggled in merely to rescue the name.
This is a useful example of how scientific dating changes historical possibility space. Before radiocarbon evidence, a Bacon attribution might have seemed chronologically open. After dating, the direct-authorship version becomes untenable.
The lesson is not that every old claim was foolish. Historical researchers worked with the evidence available to them.
The lesson is that knowledge is replaceable. New evidence should be allowed to retire explanations that once seemed plausible.
24. Has the Voynich Manuscript been decoded?
No proposed full decipherment has achieved general acceptance among relevant specialists.
That sentence needs care because “decoded” can mean several things.
A researcher might identify a candidate language. Another might propose a cipher mechanism. Another may offer readings for selected labels. Another may claim to understand the manuscript’s subject while leaving glyph values unresolved. Those are different levels of solution.
A robust full decipherment should do more than produce meaningful-looking phrases from selected passages. It should use stable rules, explain a substantial portion of the corpus, work across material not used to invent the method, reproduce independently, fit known chronology and manuscript structure, and make successful predictions.
This is developed in What a Real Voynich Decipherment Must Survive.
The absence of a generally accepted decipherment does not mean no progress has been made. Material dating, scribal analysis, transcription, statistical study, codicology and digital imaging have substantially improved the map of the problem.
Unsolved does not mean unstudied.
25. Is it a cipher?
It may be, but “cipher” remains a hypothesis class rather than an established fact.
The manuscript’s unreadable script naturally attracted cryptanalysts. If meaningful text was transformed through substitution, transposition, nomenclators, abbreviations or more complex procedures, cryptographic analysis might recover the underlying system.
However, simple cipher hypotheses face structural challenges. A straightforward monoalphabetic substitution of ordinary European prose would tend to preserve certain statistical properties that do not map cleanly onto all Voynich observations. More complicated systems can fit more behaviour but also introduce more degrees of freedom.
The key methodological danger is flexibility. If a cipher solution allows glyph values, reading direction, abbreviations and transformations to change whenever needed, meaningful output becomes easy to manufacture.
A serious cipher hypothesis should therefore become more constrained as it succeeds. Rules discovered on one sample should predict another.
The correct knowledge statement is: cryptographic explanations remain among the plausible families, but no specific cipher system has been demonstrated convincingly across the full corpus.
26. Is it a natural language?
Natural language is another major hypothesis class.
Voynichese exhibits features that invite linguistic analysis: repeated words, positional constraints, token families, page-level vocabulary differences and non-random sequence patterns.
But language-like does not mean a language has been identified.
Known languages vary enormously in morphology, orthography and information structure. Medieval writing can include abbreviations, inconsistent spelling and specialist notation. An unknown language written in an unfamiliar or transformed script could look unusual.
At the same time, constrained non-linguistic systems can imitate some language statistics. This is why isolated similarities—such as one token resembling a word in Latin, Italian, Hebrew, Turkish or another language—carry little weight without a reproducible mapping across unseen material.
A natural-language hypothesis becomes stronger when it explains several levels simultaneously: glyph behaviour, word formation, syntax, frequency, section variation and semantic alignment with independent visual or historical evidence.
That level of convergence has not yet produced general agreement.
27. Is it generated or meaningless text?
Another family of hypotheses proposes that the text may have been produced by a generative procedure rather than encoding ordinary semantic language in a straightforward way.
This idea exists in many variants.
Some imagine pseudo-text created mechanically or algorithmically. Others propose copying with controlled variation. Some suggest a hoax or prestige object designed to look meaningful. More nuanced versions allow structured but non-linguistic information.
The phrase “meaningless text” is too coarse because a generated system can still have function. It may encode categories, mnemonic cues, procedural states or visually organised information without mapping to ordinary sentence semantics.
Any generative hypothesis must explain the manuscript’s actual structure rather than merely show that a generator can produce something that looks vaguely Voynich-like. The test is multi-dimensional: token distributions, near-neighbour families, local repetition, line effects, section differences, scribal variation and relationships with imagery.
A generator that matches one statistic while failing ten others has not reproduced the system.
The responsible knowledge position is that generative models are serious comparison tools, but no simple generator has established the manuscript’s purpose or fully explained its observed behaviour.
28. Is it a hoax?
“Hoax” is one of the most common public questions and one of the least precise.
A hoax requires intent to deceive someone about what the object is. That is a historical claim about motive, not simply a statistical description of the text.
A manuscript could contain generated pseudo-text without being a hoax if the pseudo-text served a legitimate mnemonic, ritual, artistic or classificatory purpose. Conversely, a meaningful encoded text could be used deceptively in some historical context.
To establish hoax, researchers would need evidence about production circumstances, intended audience and deceptive purpose.
The material age of the manuscript also means a hoax hypothesis must be medieval or early modern in its core production story, not a simple twentieth-century fabrication by Wilfrid Voynich. The physical evidence strongly constrains modern-forgery narratives.
So the responsible answer is not “yes” or “no.” It is that deliberate deception remains one possible historical-purpose hypothesis, but “hoax” should not be used as a synonym for “we cannot read it.”
29. What language is it?
No language identification has achieved general acceptance.
Over the decades, researchers have proposed numerous languages and language families. The breadth of proposals is itself informative: flexible methods can often generate apparent matches to many languages.
A convincing identification would need more than isolated lexical resemblance. It should establish stable sound or sign correspondences, grammar or morphology, reproducible segmentation and semantic coherence across large portions of text.
It should also account for the manuscript’s unusual structural properties rather than treating them as inconvenient noise.
Historical plausibility matters too. A language needs a plausible route into the manuscript’s production environment. But historical plausibility alone cannot rescue weak textual mapping.
The correct state is therefore unresolved.
Unresolved is not failure. It is a boundary around what evidence has not yet established.
30. Who wrote it?
No named author has been established.
The manuscript has attracted associations with Roger Bacon, John Dee, Edward Kelley and many other historical figures, but fame of attribution is not evidence of authorship.
Multiple scribal hands make the question “who wrote it?” more complex than it first appears.
We need to distinguish intellectual author, compiler, scribe, illustrator, colourist, patron and later owner. One person could hold several roles, or they could be distributed across a workshop or community.
This is one reason first-principles reconstruction will matter in Article 9. Instead of beginning with a famous name, we will ask what kind of production environment could generate an object with this combination of text, imagery, specialised layout, multiple hands and material investment.
At the KNOWLEDGE stage, the correct answer remains simple: the author or authors are unknown.
31. Where was it made?
No production location has been securely established.
Yale’s current public description uses a broad Central European production framing, while research has proposed regions across Europe and beyond based on artistic, linguistic, codicological and historical clues.
The Padua branch in the eduKate research library is one sustained attempt to test a particular historical environment. It has generated many comparisons and useful observations, but it should remain a case hypothesis until it survives strong alternatives and independent evidence.
Geographic inference is exactly where Rainbolt discipline matters most. A clue is valuable when it distinguishes one place from others.
A visual motif common across much of Europe offers broad compatibility but little location resolution. A scribal convention restricted to a narrow region could be more powerful. A historical institution with the right combination of medical, astronomical and manuscript practices might be relevant, but matched institutions elsewhere must be checked.
The current responsible state is therefore a constrained but unresolved geography.
32. What was it for?
The manuscript’s purpose is unknown.
Its sections suggest systematic organisation. The combination of plants, celestial diagrams, bodies, containers and short entries encourages comparison with medical, pharmaceutical, astrological, natural-philosophical and practical compendia.
Yet none of those genre labels fully resolves the manuscript.
A medieval technical book could combine several knowledge systems because modern disciplinary boundaries did not apply in the same way. Medicine could involve astrology, plants, bodily regimes and timing. Natural philosophy could overlap with cosmology. Practical knowledge could combine recipes, diagrams and mnemonic conventions.
Another possibility is that we are misclassifying the object because its original genre is unfamiliar or hybrid.
The best route is to derive purpose from multiple independent features: page organisation, section sequence, repeated textual structures, image-text relationships, material investment and comparison with historically plausible manuscript types.
Until those converge, “purpose unknown” is more accurate than a confident genre claim.
33. What do the pictures mean?
Some pictures can be placed within broad visual categories. Exact meanings often remain uncertain.
This distinction deserves repeating because images generate an illusion of accessibility. A reader can see a plant, a human figure or a circular diagram even when the text is unreadable. That visual familiarity encourages semantic confidence.
But historical images are not photographs. They can be schematic, symbolic, copied, hybridised, mnemonic or convention-bound. A modern viewer may recognise the wrong category because the original audience used a different visual grammar.
The right question is therefore not only “What modern object does this resemble?” but “What functions did similar forms perform in relevant manuscript traditions?”
That shift from resemblance to function is crucial.
An image can be historically related without depicting the same literal object. A diagrammatic convention can migrate across domains. Decorative forms can mimic technical forms.
Our pictures article remains the best specialist route for this question. The KNOWLEDGE article preserves only the boundary: broad visual families are identifiable; many exact referents and functions are not.
34. What mathematics can tell us without translation
Unreadability does not make the manuscript mathematically silent.
We can count glyphs, tokens, lines, paragraphs, page vocabularies, repeated sequences, positional effects and transitions. We can measure entropy, conditional probabilities, clustering and similarity. We can compare sections, scribes and control corpora.
This is the central insight of Voynich | The Mathematics and Science: meaning is not the only form of information available.
Mathematics can tell us that two page groups behave differently without telling us why. It can tell us that some token transitions are more probable than others without assigning grammar. It can show that a generator fails to reproduce observed line-position effects without revealing the true generator.
This is extremely valuable because it turns some claims into falsifiable predictions.
If a decipherment implies that several surface forms are equivalent, their distributions may be expected to behave accordingly. If a proposed language model predicts certain morphology, token patterns can test it. If a generative procedure claims to explain the text, synthetic output can be compared on multiple dimensions.
Mathematics does not solve the manuscript by authority. It creates additional gates a solution must pass.
35. What science can tell us without overclaiming
Science contributes to Voynich through several modes.
Material science examines parchment, ink and pigments. Imaging can reveal details difficult to see under ordinary light. Statistics and information theory characterise text. Computational methods compare patterns. Palaeography and codicology use systematic historical methods. Botanical and iconographic comparison can be formalised more carefully than impressionistic resemblance.
The unifying scientific principle is not any one instrument.
It is replaceability.
A claim should remain open to revision when better evidence or a better model arrives.
This is especially important for an object that has resisted confident explanation. The temptation is to interpret uncertainty as an invitation to weaker standards. Science requires the opposite. The harder the ground truth is to access, the stronger the need for explicit controls, reproducible methods and confidence calibration.
The goal is not to eliminate speculation. Speculation generates hypotheses.
The goal is to prevent hypotheses from becoming facts simply because they are interesting.
36. What counts as a real advance?
Because the manuscript remains undeciphered, progress is sometimes judged too narrowly.
If there is no translation, people assume nothing important has changed.
But research advances whenever possibility space becomes more accurately constrained.
Dating the parchment is an advance.
Identifying multiple scribal hands is an advance.
Building reliable digital transcriptions is an advance.
Showing that a proposed statistical anomaly disappears under matched controls is an advance.
Demonstrating that a decipherment requires inconsistent rules is an advance.
Discovering that a visual motif is common across several regions rather than unique to one is an advance.
Each result changes the map even if the destination remains unknown.
This is why the existing 303-article research library matters. It preserves tested routes rather than only headline conclusions.
37. What the research archive currently contains
The eduKate Voynich Research Library has grown into a large experimental record rather than a conventional linear book.
Its major shelves include foundational subject pages, the “Everything eduKate Knows and Tested” sequence, comparator studies, the extensive Padua research branch and Tangential Lens experiments.
That structure has strengths and weaknesses.
The strength is provenance. Individual experiments preserve what was tested, what motivated the test, what comparisons were used and what happened. Failed routes remain visible.
The weakness is that a new reader can encounter hundreds of nodes without knowing which facts are foundational, which are hypotheses, which are methodological lessons and which represent active research branches.
The new 16-longform explanatory layer exists to solve that second problem without destroying the first.
The laboratory remains intact.
The university is being built above it.
38. A confidence ladder for Voynich claims
One practical way to keep the knowledge layer stable is to assign claims to confidence classes.
Level 1: Directly observable. The page contains a particular drawing, glyph sequence, fold, hole, stain or layout feature.
Level 2: Measured. Radiocarbon results, material composition, token frequencies, page dimensions, repeated sequences.
Level 3: Reproducible classification. Multiple analysts can identify the same broad scribal hand, Currier population or visual section under stated criteria.
Level 4: Strong inference. Evidence supports a historical or functional conclusion better than alternatives but does not directly observe it.
Level 5: Working hypothesis. A plausible explanation that generates testable predictions.
Level 6: Speculation. An idea worth noting but not yet sufficiently constrained.
Level 7: Contradicted or retired. A claim conflicts with strong evidence or fails reproducible tests.
The exact numbering is less important than the habit.
Claims should carry their epistemic state with them.
This is the seed of Article 13, EVIDENCE / PROVENANCE, and a direct requirement for eduKateAI.
39. What remains unknown
It is useful to state the major unknowns plainly.
- The exact place and circumstances of production are unknown.
- The intellectual author or authors are unknown.
- The exact role of identified scribal hands is unknown.
- The underlying writing system has not been securely identified.
- No language has been generally accepted.
- No cipher or generative mechanism has been generally accepted as a full explanation.
- The semantics of the text remain unresolved.
- Many illustrations lack secure identification.
- The manuscript’s exact purpose and original audience remain unresolved.
- The relationship among sections, production stages and scribes is only partly understood.
This list should not be read as evidence that research has failed.
It is a specification of the remaining problem.
A good unsolved problem becomes more precise over time.
40. What would change the state dramatically?
Some discoveries would alter the Voynich problem far more than others.
A securely proven parallel manuscript using the same script would be transformative.
An archival document unambiguously describing the manuscript’s production could transform provenance.
A stable decipherment that predicts unseen passages across sections would transform textual understanding.
A newly discovered key or bilingual parallel would change identifiability.
New imaging revealing erased or overwritten text could open another evidence channel.
By contrast, another visually suggestive plant match or another handful of words extracted through flexible rules is unlikely to change the state much.
This distinction is Rainbolt’s discrimination principle applied to research strategy: prioritise clues by how much they could move the map.
41. Why matched comparison is essential
Almost every strong claim about Voynich requires comparison.
The text is repetitive—compared with what?
The script is unusual—compared with which scripts?
The plants are strange—compared with which manuscript illustration traditions?
The diagram resembles a city—how many non-city diagrams share the same geometry?
Comparison creates a denominator.
Without a denominator, coincidence looks like rarity.
The challenge is that matched controls are multi-dimensional. A good comparison manuscript might need similar date, genre, region, length, scribal culture and technical purpose. No single control may satisfy all properties.
The solution is not to abandon comparison but to use several control families and state what each one controls.
Article 11, COMPARE, will make this a full reasoning engine.
42. Why negative evidence matters
Voynich theories often accumulate positive matches.
This glyph resembles that letter.
This plant resembles that species.
This diagram resembles that map.
This word resembles that historical term.
But strong theories also explain absences.
If a language hypothesis predicts a frequent grammatical marker, is it present where expected? If a regional origin theory predicts a distinctive iconographic convention, why is it missing? If a recipe genre is proposed, do entry structures show expected regularities?
Negative evidence is difficult because absence can have many causes: damage, incomplete corpus, scribal omission, genre variation or wrong expectation.
But a theory that explains only what it finds and never what it fails to find is dangerously flexible.
Article 7, ADVERSARIAL, will treat negative evidence as a first-class research tool.
43. Why the manuscript is ideal for AI research discipline
The Voynich Manuscript is a particularly demanding object for AI because the system cannot rely on a known answer key while generating explanations.
A model can retrieve dozens of proposed decipherments. It can compare plant images. It can analyse transcriptions. It can generate linguistic hypotheses. It can search historical networks.
The difficult part is not idea production.
It is epistemic bookkeeping.
Which statement came from the manuscript itself?
Which came from a materials report?
Which is a scholarly interpretation?
Which is eduKate’s own working hypothesis?
Which was generated by the AI during the current reasoning pass?
Those layers must not be allowed to merge.
This is why the KNOWLEDGE article is necessary before the MECHANISM, CASE and SCENARIO articles. eduKateAI needs a stable factual owner to return to whenever speculative branches drift.
The stable owner does not have to know everything.
It has to know what kind of knowledge each statement represents.
44. The difference between the manuscript and “Voynich”
There are really two objects called Voynich.
The first is Beinecke MS 408: parchment, ink, pigments, folds, writing, drawings and documented custody.
The second is the cultural object: the mystery, theories, documentaries, internet discussions, proposed solutions, fictional associations and reputation as the world’s unreadable book.
The cultural object affects research because it changes which questions are asked and which claims receive attention.
A dramatic decipherment attracts more attention than a careful finding about quire structure. A mysterious plant identification circulates more widely than a negative comparator result. A story about a secret society is easier to remember than a transcription caveat.
The KNOWLEDGE article therefore keeps returning to the physical manuscript as gravity.
The stories may be interesting.
The object gets the final vote.
45. What Yale’s custody means for modern research
Today the manuscript is held by the Beinecke Rare Book and Manuscript Library at Yale.
Its fragility and extraordinary research demand mean access to the physical original is restricted, but Yale provides high-resolution digital images of the entire manuscript for research.
This digital access transformed the field.
Researchers around the world can inspect pages without handling the object. Transcriptions can be aligned with images. Computational analysis can be reproduced. Visual comparisons can be checked immediately.
Digitisation also introduces a new epistemic layer. A scan is not the manuscript. Colour calibration, resolution, lighting, compression and image processing affect what appears on screen. Some questions still require physical examination or specialised imaging.
So the evidence hierarchy becomes:
- physical object;
- scientific measurement of object;
- high-resolution representation;
- transcription or derived dataset;
- analysis;
- interpretation.
Each step is useful. Each adds transformation.
46. Why the current evidence does not support one simple story
If the Voynich Manuscript were simply one ordinary language written in one ordinary substitution cipher by one person for one purpose, we might expect the problem to have yielded more easily after decades of modern analysis.
That does not prove the system is extraordinarily complex.
But the persistent resistance of simple models is itself informative.
The manuscript combines multiple visual sections, textual subpopulations, likely multiple hands, unusual token structure, positional effects and an uncertain relationship between text and image.
A successful explanation may therefore need to account for several interacting layers rather than one universal key.
Possible complexity should not become an excuse for unlimited flexibility. The correct response is decomposition.
Which features are global?
Which vary by section?
Which vary by scribe?
Which are consequences of transcription?
Which hypotheses explain several dimensions at once?
This decomposition is the bridge from KNOWLEDGE to MECHANISM.
47. The factual floor
We can now state a compact factual floor for the entire Voynich programme.
- There is a real historical illustrated parchment codex, Beinecke MS 408.
- The parchment belongs to an early fifteenth-century material timeframe.
- Materials analysis is broadly consistent with historical manuscript production.
- The script remains unidentified and the text lacks a generally accepted decipherment.
- The text is measurably structured rather than visually arbitrary.
- Distinct textual populations and multiple scribal hands have been identified at useful levels of analysis.
- The manuscript contains several visually differentiated sections.
- Some imagery belongs broadly to recognisable historical categories, while many exact meanings remain uncertain.
- Later provenance is substantially documented, while original production circumstances remain unresolved.
- No named author, language, cipher, location or purpose has achieved general acceptance as the full explanation.
Every later article in the 16-engine sequence should be allowed to explore aggressively.
None should be allowed to silently overwrite this factual floor without stronger evidence.
48. The return: what is the Voynich Manuscript?
We can now answer the question more precisely than we could at the beginning.
The Voynich Manuscript is not merely “an undeciphered book.”
It is a surviving historical information object with multiple evidence layers.
Its parchment and materials constrain chronology.
Its binding and foliation constrain construction.
Its script provides measurable structure.
Its page populations and scribal hands reveal internal variation.
Its images provide broad historical and functional clues while resisting easy identification.
Its provenance preserves part of its journey while leaving its origin uncertain.
Its century of modern research provides a second archive: a record of how humans attempt to reason when meaning is unavailable.
The manuscript is therefore both an object of study and a test of method.
We know enough to reject many careless stories.
We do not yet know enough to replace them with one generally accepted complete explanation.
That is the responsible state.
And it is not empty.
49. Codicology: the book as an engineered object
Codicology studies the manuscript as a constructed object: sheets, folds, gatherings, sewing, binding, foliation, missing leaves, repairs and the physical sequence in which a book came together. For Voynich, this matters because the codex can preserve evidence of production even where the text withholds meaning.
A medieval book was not produced by clicking “new document.” Animal skin had to be prepared, cut, ruled or otherwise organised, folded into gatherings and written. Larger foldouts required special planning. Illustrations occupied space that text could not. Mistakes, corrections and changes of plan left traces. Binding could occur after writing and could later be altered. Each physical decision creates constraints on what came before and after.
This is why page order cannot be treated naively. The order in which a modern digital viewer presents folios is the order of the surviving object, not necessarily a guarantee that every leaf stands exactly where the original production plan placed it. Missing leaves, detached folios and rebinding can complicate narrative assumptions.
For a decipherment theory, codicology acts as a gate. If a proposed reading assumes that one section logically follows another, the physical structure should be checked. If a theory requires one diagram to have been planned after a later page, construction evidence may support or weaken that sequence. If text crosses a fold in a way that implies the foldout existed before writing, that constrains production order.
Codicology rarely produces the dramatic headline “Voynich solved.” Its power lies in preventing impossible stories from surviving. In the knowledge architecture, that makes it a high-confidence owner of construction questions and a natural handoff point whenever a linguistic or historical theory begins depending on page order.
50. Folios, recto and verso: why page language matters
Voynich literature usually identifies pages by folio rather than by ordinary modern page number. A folio is a leaf. Its front side is recto and its back side verso. A reference such as f57v points to folio 57, verso side.
This convention is not pedantry. It allows researchers across different editions, image systems and discussions to refer to the same physical surface. When foldouts contain multiple panels, references can become more complex, and consistent naming becomes even more important.
A knowledge system should preserve these identifiers because they are the addresses of primary evidence. If a claim says a glyph occurs beside a certain diagram, the reader should be able to return to the exact folio. If an AI summarises a result, it should retain the evidence location rather than only the prose conclusion.
This is analogous to citing a line in a scientific dataset or a coordinate on a map. The identifier allows independent checking. Once claims lose their folio anchors, interpretation can drift away from the object.
For eduKateAI, folio references therefore belong in the provenance layer. A node such as “diagram X resembles tradition Y” should not float as an abstract statement. It should carry the exact source surface, the image version used, the comparator source and the analyst who made the comparison.
51. Missing leaves and the danger of reconstructing what is gone
The manuscript is incomplete. Missing leaves matter because they can interrupt sections, remove beginnings or endings and distort our sense of sequence.
Humans naturally repair missing structure in imagination. If a quire appears to lack one leaf, we begin picturing what might have been on it. Perhaps a missing plant completed a sequence. Perhaps a diagram once explained the next page. Perhaps a title page identified the work. Those possibilities are tempting precisely because the missing evidence cannot contradict us.
Responsible reconstruction therefore distinguishes physical inference from semantic invention. Sewing structure may establish that a leaf once existed. Offset marks or ink transfer may reveal something about adjacent surfaces. A catchword or continuation could suggest textual sequence. But the content of the missing page remains unknown unless indirect evidence constrains it.
This is an important general principle: absence creates freedom for stories. The less evidence survives, the easier it becomes to design an explanation that fits. Strong research should become more cautious in gaps, not more imaginative without labels.
The Atlas should eventually mark missing leaves as VOID nodes rather than silently bridging across them. In CivDJ terms, a VOID is not nothing. It is a known absence with structural consequences. That distinction prevents the machine from filling a historical gap with generated content merely because continuity feels preferable.
52. Marginalia and later marks
Not every mark visible in a historical manuscript necessarily belongs to the same production event. Owners, readers, librarians and later handlers can add annotations, numbers, signatures or other marks.
Voynich contains features that have attracted attention as possible later writing or marginal notation. These can be valuable because a later reader may have known something about the manuscript that is otherwise lost. But later marks can also mislead if they are treated as part of the original script.
Palaeography, ink appearance, position and historical context can help distinguish layers. A seventeenth-century ownership mark tells us something about custody, not necessarily fifteenth-century authorship. A later foliation helps us navigate the surviving codex but may not reflect original page numbering.
This is why temporal provenance should attach to annotations. A statement should specify not only what is written but which production layer it belongs to, if that can be established.
For AI, this prevents a classic retrieval error: merging evidence from different centuries into one imagined contemporaneous scene. Historical objects accumulate layers. Knowledge systems must be able to represent those layers without flattening them.
53. The problem of the alphabet
Even the question “How many letters are in the Voynich alphabet?” is less simple than it appears.
Some marks are clearly frequent and distinct. Others may be variants of the same underlying sign. Complex forms may be single glyphs, ligatures or sequences written continuously. Rare marks could be errors, decorations, abbreviations or true members of a larger inventory.
An alphabet count therefore depends on a model of graphical identity. Different transcription traditions make different decisions. Those decisions can influence statistical conclusions about character frequency, entropy and word structure.
This creates a useful hierarchy. At the lowest level we have physical strokes. Above that are perceived graphemes. Above that are transcription symbols. Above that may be hypothesised linguistic or cryptographic units. Each level introduces interpretation.
A decipherment can fail by jumping levels. It may treat a transcription symbol as though it were already a phoneme, or treat a visible space as though it were already a word boundary. KNOWLEDGE should preserve the distinction so MECHANISM can later test different unit models explicitly.
54. Spaces may be evidence without being words
The Voynich text contains visible spaces that divide strings into word-like units. Most computational work treats these as tokens because doing so is useful and because the spacing appears deliberate.
But a space is a graphical fact before it is a linguistic fact.
Historical scripts use spacing differently. Abbreviation systems can create units unlike modern words. Encipherment can preserve, remove or introduce spaces. A generative system may use spacing as part of a production rule. Even in known languages, scribal segmentation can vary.
This matters because many Voynich statistics are token-based. Word length, vocabulary size, repetition and family structure all depend partly on segmentation. A theory that changes segmentation can change the apparent statistical system.
Strong claims should therefore be tested for robustness to reasonable alternative segmentation. If an effect disappears as soon as one ambiguous space is handled differently, confidence should fall. If it survives several tokenisation choices, it becomes more interesting.
55. Lines are not neutral containers
Voynich text shows evidence that line position matters. Certain forms or token behaviours differ near line beginnings, line endings or other positional contexts.
This is a major constraint because ordinary semantic models often treat line breaks as incidental consequences of page width. If line position changes token behaviour systematically, the mechanism may include layout, scribal procedure, abbreviation, formatting conventions or line-sensitive generation.
Several explanations remain possible. A scribe could choose different abbreviations to fit available space. Formulaic text could have special opening forms. A cipher or generated system could use line-level rules. The line might represent a functional unit rather than a typographic accident.
What matters at the KNOWLEDGE stage is that positional structure deserves first-class status. Any proposed mechanism that ignores line effects is incomplete until it explains why those effects do not matter.
This is precisely the kind of observation that can distinguish candidate systems later. It is relatively dull compared with a spectacular plant match, but Rainbolt logic says dull discriminators can be more valuable than vivid similarities.
56. Paragraphs and local context
The text is organised into paragraphs or paragraph-like blocks in many parts of the manuscript. Some initial glyph forms and local vocabulary behaviours have been reported to vary with paragraph position.
Again, this provides structure before semantics. A paragraph may correspond to a topic, recipe, instruction unit, copied source segment or production batch. We do not need to know which interpretation is correct to measure how paragraphs differ.
Local context also matters for token similarity. Some near-neighbour forms cluster near each other. That can be consistent with ordinary discourse, where related words co-occur, but it can also arise from copying and modification or from local generative rules.
The key question becomes scale: how far does dependence extend? One glyph? One token? One line? One paragraph? One page? One section? Different mechanisms predict different dependency ranges.
MECHANISM will eventually treat this as a state-transition problem. KNOWLEDGE simply records that the manuscript has several nested spatial levels and that the distribution of text is not independent of those levels.
57. Labels are tempting because they look like anchors
Short strings positioned next to stars, plant parts, diagram elements or other figures are often treated as labels. This is reasonable visual language: the placement resembles labelling conventions in many manuscripts.
If they are labels, they could be unusually valuable because labels often encode names rather than full sentences. Names can connect image and text more directly than running prose.
That promise has inspired many proposed decipherments. A researcher identifies the pictured object, predicts its name in a candidate language and then maps the adjacent glyphs accordingly.
The circularity risk is obvious. If the object identification is uncertain, the predicted word is uncertain. If the glyph mapping is derived from that predicted word, the apparent successful reading is not independent confirmation.
A stronger label study asks whether groups of label-like strings show consistent statistical properties, whether the same string occurs near comparable visual objects, and whether a mapping learned from one set predicts another. Labels may indeed become anchors. The knowledge layer should preserve their promise without pretending they already have names.
58. Vocabulary is not uniform across the manuscript
Different regions of the manuscript use different token distributions. Some forms are common in one section and scarce in another. This is one of the reasons researchers can classify parts of the corpus statistically.
In an ordinary meaningful manuscript, topical vocabulary would naturally vary. A herbal section would name different things from an astronomical section. So section-specific vocabulary is compatible with semantic text.
But section differences can also arise from different scribes, production phases, source systems or generative rules. The images themselves may have influenced the writing procedure. Multiple causes can coexist.
The important distinction is between classification and explanation. Statistical clustering may reliably tell us that two page groups differ while remaining agnostic about why.
This is a useful model for AI reasoning. Classification can be high-confidence even when causal interpretation is low-confidence. The machine should not force every reliable cluster to carry a semantic label prematurely.
59. Currier categories are a doorway, not a destination
Currier A and Currier B have become so familiar in Voynich discussion that they can begin to feel like established languages or dialects. They are not.
They are historically useful labels for observable textual differences. That makes them strong descriptive categories. Their causes remain an open research problem.
The distinction becomes especially interesting when compared with scribal hands and visual sections. If one Currier population correlates with certain hands or sections, several causal stories become plausible. If boundaries cut across them, different stories gain strength.
This is exactly why relational data matters. A flat fact such as “folio X is Currier B” becomes more informative when linked to scribe classification, section, quire, token statistics and layout.
The future Atlas should therefore not represent Currier A and B as two boxes containing pages. It should represent them as one dimension across a multidimensional object. That prevents a useful classification from becoming an accidental ontology.
60. Scribal hands do not automatically equal different authors
The identification of multiple hands is often reported publicly as though several authors have been discovered. That wording outruns the evidence.
A scribe can copy someone else’s text. Several scribes can work from one exemplar. One compiler can distribute work. A master can dictate. A workshop can divide sections. Later additions can be made by different hands.
Therefore, palaeographic hand is a production role, not automatically an intellectual authorship identity.
This distinction creates testable questions. Do hands have distinct token preferences beyond ordinary variation? Do hand changes align with quire boundaries? Do they correlate with Currier type? Does illustration style change with scribal hand? Those relationships can reveal production organisation.
For the knowledge graph, SCRIBE and AUTHOR should remain separate entity roles until evidence justifies linking them. This small ontological discipline prevents a large historical narrative from being generated from one palaeographic observation.
61. Illustration style may contain several layers too
Just as writing can involve several hands, illustration and colouring can involve different production acts. Line drawing, text and colour need not have been executed by the same person or at the same moment.
This matters because colour is often used in identification. A red root or blue flower may look botanically significant. If colour was added later, copied inconsistently or applied for decorative rather than descriptive purposes, that inference changes.
Material examination and stroke relationships can sometimes help establish order. Does ink lie over pigment? Does colour respect outlines? Are different pigments used consistently across sections? These are physical questions before iconographic ones.
A robust visual analysis should therefore separate geometry, line art, colour and layout. If an identification survives without relying on possibly secondary colour, it becomes stronger.
Rainbolt reasoning again prefers clues that survive transformations. A road layout can remain diagnostic even under poor weather or camera colour. Likewise, a manuscript feature that remains distinctive after uncertain colouring is removed deserves more weight.
62. The historical context was interdisciplinary by default
Modern readers tend to classify books into disciplines: botany, medicine, astronomy, astrology, pharmacy, anatomy. Early fifteenth-century intellectual culture did not use exactly the same boundaries.
Medicine could involve celestial timing. Plants could be discussed through humoral properties and therapeutic uses. Calendrical knowledge could intersect with health. Natural philosophy and practical craft could coexist. Manuscripts could compile material from several sources rather than present one modern academic subject.
This historical difference matters because the Voynich sections can look incoherent if we insist on one modern genre. A book containing plants, zodiacal imagery, human bodies and preparations may be less anomalous within a premodern knowledge system than within a modern departmental catalogue.
But this contextual compatibility must not become a free pass. Saying “medieval knowledge was interdisciplinary” does not establish that Voynich is medical, astrological or pharmaceutical. It simply widens the set of historically plausible combinations.
This is exactly what good context should do: constrain interpretation without dictating it.
63. A manuscript can be a compilation
One underappreciated possibility is that the surviving codex may combine material from multiple sources or traditions.
Compilation was ordinary manuscript practice. A compiler could copy excerpts, rearrange material, translate, abbreviate, adapt diagrams and combine practical knowledge. The resulting book might have internal coherence without being composed from scratch as one continuous work.
This possibility matters because it changes what uniformity we should expect. Different source traditions could produce section-specific vocabulary or imagery. Scribal teams could work on different portions. Some diagrams could preserve older conventions than surrounding text.
Again, compilation is not the answer. It is a mechanism class that becomes relevant because the object shows internal heterogeneity.
The first-principles GENESIS article will eventually ask what production problem a compiler might have been solving. KNOWLEDGE keeps the possibility open while avoiding the opposite assumption that every page must derive from one authorial moment.
64. “Unknown” and “unsupported” are different states
A crucial vocabulary distinction is the difference between a question being unknown and a particular answer being unsupported.
“We do not know the language” means the language question remains unresolved.
“There is insufficient evidence that the language is X” is a statement about one candidate.
These should not collapse into “X is impossible.” A weakly supported hypothesis can remain logically possible. Conversely, the fact that many answers are possible does not make them equally probable or equally evidenced.
This distinction protects research from two opposite failures: premature certainty and indiscriminate agnosticism. We can reject poor arguments without pretending to possess the final answer. We can rank hypotheses without claiming proof.
eduKateAI should represent these states explicitly. UNKNOWN, UNSUPPORTED, CONTRADICTED and UNTESTED are different. A machine that merges them will give misleading answers even when every individual sentence sounds cautious.
65. “Possible” is not a useful confidence level by itself
Many Voynich arguments defend themselves by saying an interpretation is possible.
Possibility is a very low bar.
It may be possible that old parchment was stored for decades before use. Possible that an unusual cipher existed without surviving parallels. Possible that a plant was stylised beyond easy recognition. Possible that a rare historical transmission route connected distant traditions.
Research needs more than possibility. It needs comparative support. How plausible is this explanation relative to alternatives? What evidence does it uniquely predict? What prior historical constraints apply? What would make it less likely?
This is why Rainbolt-style discrimination improves historical reasoning. The question is not “Can I make this clue fit?” It is “Does this clue make my candidate more likely than nearby candidates?”
Possible belongs near the beginning of hypothesis generation, not at the end of argument.
66. Consensus is not proof, but it is information
Public discussion sometimes treats scholarly consensus in two unhelpful ways. One side treats consensus as unquestionable proof. The other treats it as irrelevant groupthink.
Consensus is better understood as evidence about the performance of claims under repeated expert scrutiny.
If a decipherment has been available for years and specialists across linguistics, cryptography, palaeography and manuscript studies have not been able to reproduce or accept it, that matters. It does not logically prove the decipherment false, but it lowers confidence relative to a method that independent experts can reproduce.
Likewise, consensus can change when new evidence arrives. Radiocarbon dating changed the status of older authorship claims. Better palaeographic analysis can change production models.
The correct knowledge stance is therefore neither deference nor dismissal. Ask how consensus formed, which evidence it reflects, how independent the evaluators are and what would change it.
67. Research claims need chain of custody too
Physical evidence has chain of custody. Claims should too.
Suppose a modern article states that a particular plant was identified by a famous scholar. A second article cites the first. A video cites the second. An AI retrieves all three and reports that “several sources identify” the plant.
In reality, there may be one original claim echoed three times.
Claim provenance should therefore record the earliest accessible source, later replications, critiques and revisions. Repetition should not be mistaken for independence.
This is particularly important in Voynich because many claims circulate outside peer-reviewed literature. Independent researchers can make valuable discoveries, but the route from observation to public statement may be less standardised.
The solution is not to exclude non-academic work. It is to make provenance visible enough that readers and machines can evaluate it.
68. The KNOWLEDGE Master for eduKateAI
After all these distinctions, the machine role of this article becomes clearer.
This page should operate as the Voynich KNOWLEDGE Master: the place eduKateAI returns when it needs the stable factual floor before entering speculative branches.
Its internal representation should separate OBJECT facts, MATERIAL observations, TEXT measurements, VISUAL classifications, PROVENANCE records, CLAIM histories, HYPOTHESES, CONTRADICTIONS and VOIDS.
It should know that “early fifteenth-century parchment” is not the same claim as “written in 1420.” It should know that “Currier B” is a descriptive textual class rather than a translated language. It should know that “pharmaceutical section” is a conventional visual label rather than a decoded medieval category. It should know that a proposed Padua connection belongs to a CASE branch rather than the factual floor.
Most importantly, it should provide return gravity. MECHANISM can explore. MIRROR can invert. CASE can push one hypothesis hard. SCENARIO can imagine several worlds. ADVERSARIAL can attack. Whenever those engines finish, the system returns here and asks: did anything actually change the knowledge state?
If the answer is yes, the factual floor should be updated with provenance. If the answer is no, the exploration remains an attached hypothesis rather than silently becoming fact.
The purpose of a knowledge master is not to prevent exploration. It is to make exploration return with receipts.
Primary routes from this KNOWLEDGE article
- Voynich Research Library — the full experimental Warehouse.
- Voynich Longform 01 · HUMAN — why the human mind seeks closure.
- The Manuscript Before the Mystery — physical-object discipline.
- What We Actually Know — evidence boundaries.
- What the Pictures Can and Cannot Tell Us — image evidence.
- Word Families — recurring textual structure.
- The Control Problem — how comparison disciplines unusualness claims.
- What a Real Voynich Decipherment Must Survive — the solution standard.
- Voynich | The Mathematics and Science — quantitative and scientific framing.
Authoritative external custody and material references
For the manuscript itself, the canonical external custody source is Yale University’s Beinecke Rare Book and Manuscript Library. Its public Voynich page provides current access information, provenance summary and links to the digitised manuscript. Yale also hosts the 2009 McCrone materials-analysis report.
- Yale Beinecke · The Beinecke Cipher (Voynich) Manuscript
- McCrone Associates · Materials Analysis of the Voynich Manuscript
Next: Article 3 changes engines. KNOWLEDGE tells us what is there. MECHANISM asks what the observable system is doing. We will move from the codex as an evidence object into How the Voynich Manuscript Works | The Machine We Cannot Yet Read.