Almost every computational study of the Voynich Manuscript begins with a quiet gift from the scribe.
Spaces.
The text is not one uninterrupted river of glyphs.
It contains visible gaps.
Those gaps divide the writing into manageable strings.
We can count them.
Sort them.
Find frequent forms.
Measure word length.
Build word families.
Compare Currier A and B.
Search for keywords.
It is so convenient that our language immediately becomes confident.
We call the strings words.
Then a subtle shift occurs.
A visible gap becomes a lexical boundary.
A space-delimited string becomes one unit of language.
A recurring string becomes a recurring word.
And before long, a theory has inherited the entire architecture of ordinary written language without proving that the Voynich writer used spaces for the same job.
A space is an observation. A word boundary is an interpretation of that observation.
The distinction is small in typography.
It is enormous in decipherment.
Quick Read
One-sentence answer: Voynich writing clearly uses visible spacing to divide much of its linear text into recurring token-like units, and those units behave systematically enough to support powerful statistical analysis, but no accepted decipherment proves that every space corresponds exactly to a lexical word boundary rather than to a syllable group, morpheme bundle, cipher group, abbreviation unit, scribal rhythm unit or another documentary convention.
- Visible spaces in the manuscript are real and analytically useful.
- Calling the resulting strings “words” is conventional shorthand, not a deciphered fact.
- Many Voynich studies therefore use “word”, “token” or “word-like unit” with different levels of caution.
- Space-delimited units have strong internal structure and recurring families.
- Changing one uncertain space can create or destroy a token, changing vocabulary counts and word-family statistics.
- Word length depends on both space boundaries and glyph segmentation.
- Line-position behaviour can interact with spacing and token shape.
- Labels complicate the picture because they often occur as isolated strings where the physical layout itself creates obvious boundaries.
- Circular text complicates it further because spatial separation and linear order are different problems.
- Natural-language writing can use spaces imperfectly or at units larger or smaller than modern lexical words.
- Historical abbreviation can make one visible token represent a longer linguistic sequence.
- Cipher systems can use spaces to separate code groups rather than plaintext words—or can preserve plaintext word spacing.
- A structured notation can use gaps as field separators.
- The correct model should explain why the observed boundary system produces the manuscript’s word lengths, repetition, morphology-like families, line effects and label behaviour together.
Spaces give us a segmentation.
They do not yet tell us the linguistic level of that segmentation.
Why Spaces Feel More Certain Than They Are
Modern readers are trained to treat white space as grammar.
The cat sleeps contains three words because modern English orthography tells us so.
But writing systems do not all segment language the same way.
Some historical texts have inconsistent spacing.
Some scripts traditionally write continuously.
Some orthographies attach clitics that another language writes separately.
Compounds can be written as one word, two words or hyphenated depending on convention.
Therefore the existence of spaces is stronger evidence than their exact semantic interpretation.
The Voynich scribe clearly distinguished some units visually.
What those units correspond to remains part of the problem.
“Token” Is the Safer First Word
A token is simply an analytical unit produced by a segmentation rule.
If we split Voynich linear text at visible spaces, we obtain space-delimited tokens.
This statement carries almost no semantic assumption.
Then we can ask whether those tokens behave like words.
- Do they recur?
- Do they have constrained internal structure?
- Do common tokens distribute like function words?
- Do rare tokens cluster like content words?
- Do token boundaries align with grammatical behaviour?
- Do labels look like a noun-like subset?
The token gives us a neutral starting point.
“Word” becomes a hypothesis that can earn strength.
The Spaces Are Useful Because the Resulting Units Are Structured
If Voynich spaces were arbitrary decoration, splitting at them might produce chaotic strings.
Instead, the resulting tokens show strong regularities.
- Recurring forms are common.
- Token lengths occupy a constrained range.
- Glyph positions inside tokens are strongly restricted.
- Some openings and endings form dense families.
- Currier A and B differ in token distributions.
- Line beginnings and endings affect token composition.
This tells us that the spaces are not useless.
They carve the text at boundaries that correspond to genuine structural organisation.
The unresolved question is the level.
Lexical word?
Morphological group?
Code group?
Production unit?
The boundary works.
Its interpretation remains open.
One Space Can Change an Entire Vocabulary Table
Imagine the visible sequence:
ABCDE FGH
One transcriber records two tokens.
Another decides the gap is scribal and records:
ABCDEFGH
Now:
- token count changes;
- average token length changes;
- ABCDE may become a frequent word in one corpus and disappear in another;
- prefix and suffix statistics change;
- edit-distance families change;
- entropy around the boundary changes.
This is why uncertain spaces are not clerical trivia.
They can alter the very object called “vocabulary”.
Not Every Gap Is Equally Clear
Handwritten spacing is continuous.
There is no invisible ruler between every pair of glyphs.
Some gaps are obviously large.
Some are obviously within connected writing.
Some are borderline.
That creates inter-transcriber disagreement.
A robust study should therefore know whether its central result depends on high-confidence spaces or on a cluster of ambiguous ones.
Where multiple transliterations disagree, the disagreement is itself useful metadata.
The representation can remain uncertain rather than forcing every gap into yes/no certainty.
Word Length Is Not a Raw Property of the Manuscript
Researchers often compare Voynich word lengths with natural languages.
This can be useful.
But word length depends on two hidden decisions.
- Where does one token end?
- How many characters does each visible form contain?
If ckh is one grapheme rather than three, a five-EVA-character token may become three underlying characters.
If one visible gap is not a lexical boundary, two short tokens may become one longer word.
Therefore “Voynich words are unusually short” or “Voynich words have this average length” is always conditional on a segmentation and alphabet.
The statistic can still be correct.
Its ontological depth should be stated honestly.
Natural Languages Do Not Agree on What a “Word” Looks Like
Even familiar languages make word boundaries complicated.
English writes ice cream as two words but blackbird as one.
Languages differ in whether pronouns or particles attach to verbs.
Historical spelling conventions can shift over time.
Agglutinative languages can pack material that English would express with several words into one long form.
Therefore a Voynich space system that does not correspond exactly to modern English wordhood would not be surprising.
The stronger question is whether the segmentation corresponds consistently to some linguistic or documentary unit.
Spaces Could Separate Morpheme Bundles
Suppose Voynich writes units smaller than lexical words.
A stem might be separated from a grammatical particle.
Or a long underlying word might be broken into pronounceable or writable groups.
This could explain why many tokens are highly structured yet difficult to map onto expected word-frequency distributions.
A morpheme-group model makes predictions.
- Certain token sequences should recur together.
- Boundaries should align with stable grammatical combinations.
- Some adjacent tokens should behave more tightly than ordinary word pairs.
Without such evidence, “morpheme groups” is merely another possible label.
The mechanism needs to generate the observed boundary statistics.
Spaces Could Separate Syllabic or Phonological Groups
Another possibility is that visible units correspond to chunks of sound rather than lexical words.
Some writing traditions and pedagogical notations can segment language by syllable or pronunciation unit.
If Voynich tokens are syllabic groups, then repeated short forms need not be function words.
Word-frequency comparisons with ordinary orthographic corpora would become mismatched.
A syllabic-group theory should eventually reconstruct longer underlying lexical words from stable token sequences.
If no such higher-level regularity emerges, the idea remains speculative.
Spaces Could Separate Cipher Groups
Ciphertext does not have to preserve plaintext spaces.
Some systems remove them.
Some preserve them.
Some introduce grouping for the convenience of the encipherer or reader.
Therefore Voynich spaces could separate encoded groups whose boundaries only partly correspond to underlying words.
A group-based cipher can explain why visible “words” have tight structural constraints.
But it must also explain why the same groups exhibit local vocabulary, label differences, Currier states and line effects.
Again, compatibility is cheap.
A real cipher model should generate the spacing rule.
Spaces Could Separate Fields in a Technical Notation
Not every structured writing system is ordinary prose.
A record can contain fields.
Name.
Quantity.
Operation.
Timing.
Category.
Spaces can separate those fields.
Under such a model, frequent tokens might be field codes rather than function words.
Quire 20 is especially relevant because its repeated star-marked records create a possible environment for testing repeated field order.
If token positions within records are highly regular, a field-based notation becomes more plausible.
If not, ordinary prose or another model may fit better.
Abbreviation Can Hide More Language Inside One Token
Suppose a visible token is genuinely one orthographic word.
It still need not correspond character-for-character to one ordinary spelled word.
Medieval abbreviations can compress long sequences into short forms.
This means the space boundary may be linguistically meaningful even while character counts are misleading.
A short Voynich token could expand into a longer word.
A common visible ending could encode several letters.
This is one reason low visible word length and low character entropy do not automatically rule out natural language.
But abbreviation must be systematic.
One sign cannot expand differently whenever the translator needs a new word.
Line Position Complicates Word Boundaries
The line article showed that Voynich token forms change at physical boundaries.
That creates a difficult possibility.
A line-final space may not be equivalent to an internal space.
A token could be modified because it is approaching the margin.
A new line may reset a writing or encoding state.
Repeated sequences rarely crossing line boundaries strengthens the possibility that the line imposes an additional segmentation layer.
Therefore Voynich may contain nested boundaries:
- glyph;
- token;
- line;
- paragraph;
- record;
- page.
A decipherment should tell us which are linguistic and which are production boundaries.
Paragraph Boundaries Are More Explicit Than Ordinary Spaces
Paragraphs create a stronger visual reset.
They often begin with distinctive forms or gallows-rich openings.
This suggests hierarchy.
Not all white space has equal status.
A small token gap may be lexical.
A line break may be a production boundary.
A paragraph break may mark a new semantic or record unit.
One of the future challenges is to reconstruct this hierarchy rather than treating every separator as the same kind of boundary.
Labels Are a Boundary Extreme
A label may stand alone beside a figure or object.
There is no ambiguity about where the visible unit stops because surrounding empty space isolates it.
But that does not tell us whether the label itself is one lexical word.
It could be a name.
A code.
A number.
An abbreviation.
A short phrase compressed without internal spacing.
Labels therefore show that visual isolation is not the same as lexical identification.
The next article in this batch takes that population separately because its vocabulary distribution differs strongly from running text.
Circular Text Makes “Word Order” and “Word Boundary” Separate Problems
A ring of Voynich text may contain obvious gaps between token-like units.
Those gaps can define tokens.
Yet the ring may have no obvious first token.
This separates two assumptions that linear prose lets us blur:
- where units are separated;
- in what order the units are read.
A transcription must cut the circle somewhere.
That creates a modern sequence start.
The spaces may be original.
The first token may not be.
This is why geometry must remain attached to tokenisation.
Spaces Matter for Zipf-Like and Frequency Analysis
Word-frequency distributions are often used to compare unknown texts with language.
But the distribution depends on tokenisation.
Merge two common tokens and one frequent word disappears.
Split a common token and two new frequent units appear.
A spacing system that segments morphemes rather than words can therefore produce a frequency curve that differs from ordinary lexical corpora even if the underlying message is natural language.
This does not make frequency analysis useless.
It tells us that frequency claims should be conditional on the segmentation model.
The better question is not merely:
Does the word-frequency curve look language-like?
It is:
Under what unit definition does the distribution become most coherent with the rest of the writing system?
Spaces Matter for Morphology
Suppose two Voynich tokens repeatedly occur next to each other.
A modern reader treats them as two words.
But perhaps they are one lexical word split into two morphographic groups.
Now what looked like syntax becomes morphology.
The reverse is possible too.
A long token may contain two lexical units written together.
What looked like morphology becomes syntax.
This is why a proposed root-and-affix grammar should test whether its boundaries correspond to visible spaces or cross them systematically.
A true language model should eventually recover a stable hierarchy:
grapheme → morpheme → lexical word → phrase → clause.
We do not yet know where Voynich spaces sit on that ladder.
Spaces Matter for Word Families
The previous article mapped dense families of near-related tokens.
But family membership depends on boundaries.
If AB CDE becomes ABCDE, edit-distance relationships change completely.
A prefix-like short token might actually be a detached prefix.
Or it might be a separate function word.
Those explanations predict different co-occurrence patterns.
If a short token is a detached affix, it should attach to a constrained family of following forms.
If it is an independent word, its syntactic distribution may be broader.
Spacing therefore becomes a testable part of morphology rather than a fixed precondition.
Spaces Matter for Entropy
Character entropy can be calculated with or without spaces included as symbols.
More importantly, token boundaries change the context used to interpret character placement.
A glyph that appears only at token starts becomes highly predictable if the token boundaries are correctly identified.
If the boundaries are wrong, that positional rule weakens or shifts.
This means the strong positional entropy of Voynich can help evaluate segmentation.
A better boundary model should make coherent positional constraints emerge without requiring arbitrary exceptions.
But optimisation must be controlled.
If we change boundaries solely to minimise entropy, we can manufacture an artificially tidy system.
Palaeography and visible spacing must remain independent witnesses.
The Scribe May Know the Boundary Even if We Do Not Know Its Linguistic Name
This is perhaps the most useful current position.
The writer repeatedly leaves gaps in systematic places.
The units on either side behave differently from arbitrary substrings.
Therefore the scribe was not spacing randomly.
The boundary mattered operationally.
We simply do not yet know whether the writer would have understood the separated unit as:
- a word;
- a name;
- a syllabic block;
- a code group;
- a field;
- another specialised unit.
The boundary can be real before our category for the bounded thing is correct.
What Survives the Boundary Work
- Voynich linear writing contains deliberate-looking visible spacing.
- Space-delimited units support strong, reproducible structural analysis.
- The resulting tokens have non-random internal composition.
- Token boundaries interact with line position, Currier regime and word families.
- Some spacing decisions are less visually certain than others.
- Multiple transcriptions can preserve disagreement about difficult boundaries.
- Word-length and vocabulary statistics depend on the space model.
- Natural language is compatible with non-modern word segmentation.
- Abbreviation can compress large linguistic units inside short visible tokens.
- Cipher and notation systems can also create meaningful group boundaries.
- No accepted solution proves that every visible space equals one lexical word boundary.
What Does Not Survive as Established Knowledge
- Every Voynich token is proven to be one lexical word.
- Every visible gap is proven semantically equivalent.
- Short tokens are proven function words.
- Long tokens are proven content words.
- Spaces are proven plaintext word separators.
- Spaces are proven cipher-group separators.
- Word-frequency curves can identify the language without a segmentation model.
- Modern EVA tokenisation is identical to the original writer’s linguistic ontology.
What we have is a very useful segmentation convention rooted in real manuscript spacing.
What we do not yet have is its decoded linguistic level.
A Better Boundary Analysis
- Start from visible spacing, not assumed words.
- Mark uncertain gaps rather than forcing certainty.
- Compare major transcriptions.
- Separate glyph segmentation from token segmentation.
- Measure how boundary choices affect word length and frequency.
- Control for line and paragraph position.
- Separate labels, running text and circular writing.
- Test whether recurring adjacent token sequences behave like higher-level units.
- Require a language, cipher or notation model to generate the observed spaces naturally.
What Would Count as a Real Word-Boundary Breakthrough?
Imagine a decipherment recovers a stable underlying language.
Most visible spaces align with lexical boundaries.
A smaller, predictable class marks clitic or morpheme boundaries.
The rule explains ambiguous spaces and predicts unseen lines.
That would transform tokenisation into linguistic segmentation.
Or imagine an encoding procedure demonstrates that one plaintext word is regularly transformed into two or three visible groups separated by spaces.
The procedure reconstructs the same grouping on unseen text.
Then the spaces become cryptographic group boundaries.
The breakthrough is not deciding that spaces matter.
They clearly do.
It is identifying what level of the information system they separate.
Primary School: Spaces Are Rules We Learn
Write:
ice cream
Then:
blackbird
Ask why one has a space and the other does not.
The answer is not simply meaning.
It is spelling convention.
This teaches that spaces are part of a writing system rather than natural cracks between ideas.
Lower Secondary: Change One Boundary
Give students ten invented strings separated by spaces.
Then move one space.
Ask them to recalculate:
- number of words;
- average word length;
- most frequent form;
- word families.
One tiny representational change alters several conclusions.
The lesson is immediate:
measurement depends on what you decided to count.
Upper Secondary: Word, Morpheme or Code Group?
Give students one segmented invented message.
Ask them to build three models:
- each token is a word;
- each token is a morpheme group;
- each token is a cipher group.
Then ask what each predicts about repeated adjacent pairs, token length and translation.
Students learn that the same visible segmentation can sit at different hidden levels.
JC and Adult Readers: Tokenisation Is Model Selection
At a higher level, tokenisation is not preprocessing outside the research question.
It is a model of where units exist.
Different tokenisations induce different frequency distributions, morphology, entropy and syntactic candidates.
The right tokenisation should therefore gain support from independent behaviour.
It should make grammatical relationships more systematic.
Or make an encoding procedure simpler.
Or improve prediction on unseen text.
The strongest segmentation is not the one that produces the prettiest vocabulary table.
It is the one that makes several independent layers of the manuscript more coherent at once.
A Parent and Teacher Guide
- Distinguish visible gap from lexical word.
- Use “token” when meaning is unknown.
- Show how one moved space changes statistics.
- Remember that languages segment meaning differently.
- Keep glyph and word segmentation separate.
- Ask whether boundaries change at lines, paragraphs or labels.
- Require the final model to explain why the writer put spaces where they are.
The wider lesson is simple:
a boundary can be real even when we have not yet identified the thing it bounds.
Reader Checklist: Before You Call a Voynich Token a Word
- Is the space visually clear?
- Do major transcriptions agree on it?
- Are you using “word” descriptively or linguistically?
- Could the boundary separate morpheme or syllable groups?
- Could it separate cipher or code groups?
- Does line position alter the token?
- Does glyph segmentation alter its length?
- Does the token behave consistently across Currier regimes?
- Do repeated adjacent tokens suggest a larger hidden unit?
- Do labels use the same spacing conventions?
- Does the proposed language or encoding model generate the spaces naturally?
- What would change if one ambiguous space moved?
Frequently Asked Questions
Does the Voynich Manuscript have spaces?
Yes. Much of the linear writing contains visible gaps that divide the text into recurring token-like units.
Are those units real words?
They may be, and many linguistic analyses treat them as word-like units. No accepted decipherment has yet established that every visible space corresponds exactly to a lexical word boundary.
Why use the term token?
Because “token” names the segmented string without assuming what linguistic or encoding unit it represents.
Are all spaces equally clear?
No. Many are obvious, but handwritten spacing can be ambiguous. Different transcriptions sometimes make different boundary decisions.
Could spaces separate cipher groups?
Yes. Encoded systems can preserve, remove or introduce grouping independently of plaintext word boundaries. A specific cipher theory would need to explain the observed grouping consistently.
Could one Voynich token represent more than one word?
In principle. Heavy abbreviation or a code system could compress multiple linguistic elements into one visible unit. No accepted model demonstrates this globally.
Why do spaces matter for entropy?
Token boundaries define positional contexts. Changing a boundary alters which glyphs are treated as token-initial, token-final and internal, which changes conditional statistics.
What is the strongest current conclusion?
The Voynich scribe used visible spacing systematically enough to create meaningful structural units for analysis. The exact linguistic or encoding level represented by those spaces remains unresolved.
Related eduKateSG Reading
- Voynich | Everything eduKate Knows and Tested | EVA, Transcription and the Segmentation Problem
- Voynich | Everything eduKate Knows and Tested | Word Families
- Voynich | Everything eduKate Knows and Tested | The Line as a Unit
- Voynich | Everything eduKate Knows and Tested | Why Voynichese Is So Predictable
Research and Further Reading
- Claire Bowern & Luke Lindemann — The Linguistics of the Voynich Manuscript
- René Zandbergen — Transliteration of the Voynich Manuscript Text
- René Zandbergen — Special Topics: Transliteration and Locus Structure
- René Zandbergen — The Voynich Writing System
- Luke Lindemann & Claire Bowern — Character Entropy in Modern and Historical Texts
The Final Idea
The Voynich spaces are generous.
They let us count before we can read.
They let us discover families.
Compare sections.
Measure predictability.
Find local vocabulary.
But they also tempt us to smuggle a solved concept into an unsolved script.
“Word”.
Perhaps the term is correct.
Perhaps it is only approximately correct.
Perhaps the spaces separate a different level entirely.
The strongest current position is therefore precise without being timid:
The manuscript clearly contains a deliberate boundary system. We can study the units it creates long before we know what the original writer called those units.
Continue Through the Voynich Research Map
This article is one specialist node in eduKateSG’s larger Voynich research library. Return to the canonical master to see the physical, visual, textual and historical evidence in one continuous argument.
Structure is evidence; structure is not translation. The master keeps resemblance, statistical structure, historical possibility and decipherment claims at separate evidentiary levels.
Next Best Routes
- EVA, Transcription and the Segmentation Problem — move back from token boundaries to the representation decisions underneath them.
- The Minim Strings — compare uncertainty inside a visible token with uncertainty between tokens.
- Syntax Before Semantics — test whether changing boundaries strengthens or weakens neighbouring-token structure.
Deep bridge: Compression and Reconstruction gives the right external question: does the proposed grouping preserve enough structure to reconstruct one intended source rather than many convenient alternatives?