If Voynichese is language, where are the vowels?
It sounds like a simple alphabet question.
Find the characters that behave like a, e, i, o, u.
Separate them from the consonants.
Then test whether Voynich words begin to look phonotactically ordinary.
Except Voynich refuses the first step.
In many alphabetic languages, a simple statistical model can often divide letters into two broad states that resemble vowels and consonants because vowels and consonants prefer different neighbours.
Apply the same idea to Voynich and the clean split does not reliably appear.
One famous 2011 experiment by Sravana Reddy and Kevin Knight produced something stranger: one hidden state was dominated by the final character of words, while the other state generated most of the remaining characters—as though “vowel-ness” lived at the ends of tokens rather than alternating through them in the familiar way.
Later researchers have found that the result is sensitive to modelling choices and difficult to reproduce cleanly. René Zandbergen’s more recent summary explicitly cautions that the separation into vowel-like and consonant-like states is not nearly as clear as it is in known plaintexts.
This does not prove Voynich has no vowels.
It tells us something more important:
the visible Voynich character is not yet entitled to be treated as one ordinary alphabetic letter.
A glyph may be a consonant.
A vowel.
A syllable.
A ligature.
An abbreviation.
A code group.
A positional variant.
Or one component in a larger written unit.
The vowel problem is therefore not merely “which glyph is E?”
It is a model-selection problem about the level at which phonological information, if any, is represented.
Quick Read
One-sentence answer: no accepted Voynich analysis has recovered a stable manuscript-wide vowel/consonant system comparable to ordinary alphabetic plaintext; a widely cited two-state HMM produced an unusual word-final-versus-non-final split, later work has found the separation sensitive and unclear, and the result remains compatible with omitted vowels, syllabic writing, abbreviation, compound glyphs, unusual orthography, cipher or non-linguistic notation.
- Vowels and consonants often have different neighbour statistics in alphabetic languages.
- Two-state hidden Markov models can sometimes recover those classes without knowing the language.
- Reddy & Knight applied such a model to Voynich characters in 2011.
- Their reported Voynich result did not resemble a clean ordinary vowel/consonant division.
- Instead, one state was strongly associated with word-final characters and another with most other positions.
- Zandbergen reports that the historical result depends strongly on initial conditions and is difficult to reproduce robustly.
- Later individual HMM experiments have claimed more vowel/consonant-like separations, but no consensus mapping has emerged.
- A failed clean split does not prove the absence of vowels.
- An abjad-like system may omit most vowels.
- A syllabary may encode consonant-plus-vowel combinations inside one sign.
- Abbreviation can omit predictable vowel material.
- A cipher may distribute one phonological category across several visible groups.
- Analytical transliteration can split one functional sign into components and thereby distort vowel tests.
- Synthetic transliteration can merge several functional letters and distort the same tests in the opposite direction.
- The q-series, benches, gallows and minim endings show strong positional architecture that can dominate simple phonological clustering.
- A real vowel model should improve unseen-word legality, morphology and source-language reconstruction under fixed rules.
The correct public conclusion is therefore modest.
Vowels remain possible.
The visible script has not yet told us where they live.
Why Vowels Are Statistically Discoverable in the First Place
Vowels are not simply letters whose shapes we recognise.
They occupy structural roles.
In many alphabetic languages, consonants and vowels alternate enough that their neighbour distributions differ.
A consonant often has a high probability of being followed by a vowel.
A vowel may be followed by many consonants.
Some clusters are legal.
Others are rare.
This creates statistical separability.
A two-state model can exploit that structure even without being told which symbols are vowels.
In a suitable alphabetic text, one hidden state often becomes vowel-heavy and the other consonant-heavy.
That makes the method attractive for Voynich.
If the manuscript is written in a simple substitution alphabet, the vowel/consonant alternation might survive even when every visible shape is unknown.
If it does not survive, one of our assumptions may be wrong.
- The language may have unusual phonotactics.
- The visible units may not be letters.
- Vowels may be omitted.
- Several letters may be compressed into one sign.
- One letter may expand into several visible signs.
- The text may be encoded.
- The text may not represent speech directly.
The test is therefore useful even when it fails.
The Reddy & Knight Result Is Famous Because It Fails Strangely
Reddy and Knight’s 2011 paper remains one of the important computational studies of the manuscript because it treated several basic language hypotheses as measurable questions.
Among their experiments was a two-state bigram hidden Markov model over Voynich characters.
In ordinary alphabetic language, such a model can separate vowels and consonants surprisingly well.
Voynich behaved differently.
The striking reported split put the final character of words into one state and almost everything else into another.
That is not what a normal alphabetic vowel system should look like.
But it is highly consistent with something we already know from several other directions:
word-edge position is unusually important in Voynichese.
Minim strings concentrate toward endings.
Some character families are strongly initial.
Line beginnings and endings affect token form.
The HMM may therefore have discovered positional architecture rather than phonological class.
That possibility is more valuable than forcing the two states into “vowel” and “consonant” labels.
The Result Is Not Robust Enough to Become a Voynich Law
One attractive computational result should not become scripture.
Zandbergen reports that the Reddy/Knight experiments were sensitive to initial conditions and difficult to reproduce in a stable way.
He also notes later work in which a two-state HMM was interpreted more optimistically as separating vowels and consonants.
His own broader investigation remains cautious: the separation does not appear nearly as clean as it does in known plaintext.
This disagreement is exactly what a healthy research state should look like.
The correct conclusion is not “the model proved vowels are word-final”.
Nor “another run proved ordinary vowels”.
It is:
simple two-class phonological clustering has not yet produced a stable, independently accepted Voynich vowel inventory.
The uncertainty belongs in the result.
Perhaps Voynich Is an Abjad-Like System
Some writing systems represent consonants more explicitly than vowels.
Arabic and Hebrew traditions provide familiar examples of consonant-centred writing, although their actual historical systems are richer than the simple label “no vowels”.
If Voynich omits many vowels, several observations become less surprising.
- Visible words can be short.
- Vowel/consonant clustering becomes weak.
- Consonant skeletons can produce dense word families.
- Several underlying words may collapse toward similar visible forms.
Reddy and Knight noted that Voynich word-length behaviour shares similarities with devowelled English, Arabic and Pinyin-style comparisons under certain representations.
This keeps an omitted-vowel model alive.
But an abjad hypothesis has obligations.
After supplying plausible vowels, the resulting language should become more coherent.
Consonant skeletons should map repeatedly to related lexical families.
Vowel restoration should be constrained by grammar rather than invented independently for every token.
“Maybe vowels are omitted” is a mechanism class.
It is not yet a decipherment.
Perhaps Voynich Is Syllabic
Another possibility is that a visible sign carries both consonant and vowel information.
A syllabary does not need separate vowel letters in the same way an alphabet does.
One symbol can represent ka.
Another ki.
Another ku.
In a more compositional system, a base sign may combine with a modifier to change the vowel.
Voynich’s compound-looking benches, gallows and repeated-stroke families naturally make this possibility worth testing.
For example, a bench frame could in principle be a consonantal base while an internal component changes vocalisation.
Or a minim count could distinguish syllabic values.
None of those assignments is established.
The important point is methodological:
a failure to find separate vowel characters is expected if vowels are encoded inside larger units.
The alphabet model and the syllabic model therefore make different segmentation predictions.
Perhaps the Vowels Are in the Components We Currently Merge
The reverse segmentation problem is equally important.
Suppose one synthetic Voynich glyph actually combines consonant plus vowel.
Then treating the whole compound as one character hides the vowel distinction.
Pedestalled gallows are a natural example of the general issue.
Are they atomic signs?
Or a bench-like frame plus an inserted gallows?
If the components carry independent information, phonological clustering should be performed on the components.
If the whole compound is functional, splitting it manufactures character transitions that never existed for the historical reader.
The vowel question therefore depends on the Alphabet Problem.
We cannot confidently classify characters by phonological role before we know what a character is.
Perhaps the Vowels Are in Units We Currently Split
Analytical transliteration creates the opposite risk.
EVA deliberately decomposes certain visually composite forms so researchers can inspect their components.
That is useful.
But if an EVA sequence functions historically as one syllabic or abbreviatory unit, the individual components should not be expected to divide neatly into vowels and consonants.
The minim family makes this vivid.
in, iin, iiin can be analysed as repeated strokes plus a terminal element.
Other historical transliterations have treated some such strings more synthetically.
If the entire cluster represents one abbreviation, syllable or ending, an HMM over the internal strokes is analysing pen construction rather than phonology.
Representation changes the question before it changes the answer.
Abbreviation Can Remove Vowels Without Being an Abjad
Medieval abbreviation provides another route.
A word can be written with enough of its beginning, ending or conventional shape for a trained reader to restore omitted material.
The omitted letters may include vowels.
Contractions can preserve only selected letters.
Special signs can represent common endings containing both vowels and consonants.
A heavily abbreviated Latin, Romance or technical text can therefore fail a naive character-level vowel test while still encoding ordinary vocalised language underneath.
This is why the Abbreviation or Alphabet article treated scribal compression as a system-level hypothesis.
If abbreviation explains the vowel problem, expansion should reveal stronger phonological and morphological structure.
The recovered vowels cannot be inserted freely.
Historical abbreviation rules must constrain them.
A Cipher Can Destroy a Visible Vowel Class
Ciphertext does not have to preserve the vowel/consonant classes of plaintext.
A simple monoalphabetic substitution largely does.
A verbose cipher may not.
A homophonic cipher can distribute one plaintext vowel across several ciphertext signs.
A code can map syllables or common word fragments into groups.
Nulls can disrupt alternation.
Abbreviation before encryption can remove vowels first.
The 2025 Naibbe work is relevant here because it demonstrates constructively that hand-executable historical-style encryption can transform meaningful plaintext while reproducing several Voynich-like statistics.
Naibbe is not a solution.
It is a reminder that “cipher” should not be represented by simple substitution alone.
A vowel detector that works on plaintext may fail on a realistic encoded layer.
A Generator May Produce Vowel-Like Classes Without Vowels
The reverse danger also exists.
A structured pseudo-text generator can create two character classes that statistically resemble vowels and consonants even if no sound is encoded.
Imagine a token template:
opening → core → optional middle → ending.
Characters allowed in the core will have different neighbour distributions from characters allowed at the edges.
A two-state clustering algorithm may interpret those distributional roles as phonological classes.
This is why even a successful vowel/consonant-looking split would not prove speech.
The classes must then participate in a source-language system.
Phonotactics.
Morphology.
Lexical recurrence.
Grammar.
Vowel-shaped statistics are not automatically vowels.
The q-Series Is a Warning Against Naive Phonology
EVA q is almost always followed by o.
That pair is so strong that it immediately resembles the Latin-script qu habit.
But resemblance is dangerous.
If q and o were simply /k/ + /w/ or another familiar sequence, their wider positional and Currier behaviour should support that mapping.
Instead, q behaves like a highly constrained token-initial gate and varies strongly by text regime and document role.
It is nearly absent from labels.
That could still have a phonological explanation.
It could also be grammatical, abbreviatory, cryptographic or structural.
Calling EVA o a vowel because it resembles our letter o would be an even more basic error.
EVA is a neutral transliteration label.
It is not a sound assignment.
Word-Final Structure May Be Hiding the Vowel Signal
The unusual HMM behaviour points toward a deeper possibility.
Voynich tokens may have such strong positional grammar that it dominates phonological statistics.
Suppose token endings are produced from a restricted family.
Those ending characters will share neighbour distributions regardless of whether they are vowels, suffixes or scribal terminal forms.
An unsupervised two-state model may discover the edge class first because it is statistically stronger than the vowel class.
This suggests a better workflow.
- Model token position explicitly.
- Remove or condition on edge effects.
- Then ask whether residual character behaviour separates into phonological classes.
If vowel-like separation appears only after positional structure is controlled, that would be highly informative.
If it still does not, alphabetic plaintext becomes a more difficult model.
Currier A and B Could Hide Different Vowel Realisations
Whole-manuscript clustering mixes different text regimes.
Currier A and B differ in common character groups and token families.
RZ’s newer classification also places many astronomical and cosmological pages in an intermediate C-like state.
If these regimes represent dialects, scribal conventions or encoding states, vowels may be realised differently across them.
A global HMM can blur a real within-regime distinction.
Therefore vowel analysis should be repeated:
- within A;
- within B;
- within intermediate/C material;
- within labels;
- within matched document roles.
If one stable consonant/vowel partition transfers across all regimes, it becomes a strong manuscript-wide result.
If each regime needs a completely different partition, the theory needs an explanation for that difference.
Labels Are a Particularly Hard Vowel Test
Labels have different distributions from prose.
If they are proper names, they may preserve phonology strongly.
If they are coordinates or identifiers, they may not.
A proposed vowel system derived from prose should therefore be applied to labels without retuning.
Do candidate labels become pronounceable under the same rules?
Do repeated name-like patterns emerge?
Or does the vowel model collapse completely?
Labels are useful because their document role is different enough to test transfer.
A vowel system that exists only in the corpus used to infer it is weak.
The Crib Problem and the Vowel Problem Are Coupled
A secure proper-name crib could transform vowel research.
Suppose one star or plant name were identified independently.
Its historical pronunciation would constrain which Voynich units must encode vowels.
Those assignments could then be tested across the manuscript.
The reverse is dangerous.
Assume which signs are vowels.
Insert whatever vowels are needed.
Find a name that fits.
Use the name to validate the vowel assignment.
That is circular.
A real crib should reduce the vowel possibilities before the target name is chosen.
Word Length Gives the Vowel Hypothesis Another Test
If vowels are omitted, visible words should be shorter than their fully vocalised source forms.
They may also have a narrower length distribution.
That makes the Word-Length Problem, the fourth article in this batch, directly relevant.
Reddy and Knight observed that some Voynich word-length behaviour resembles devowelled English or naturally consonant-heavy writing systems more than ordinary English spelling.
If a proposed vowel-restoration model expands Voynich words into a more normal historical length distribution while preserving grammar and lexicon, that is positive evidence.
If inserted vowels merely make arbitrary target-language words possible, it is not.
Vowel restoration must improve several independent statistics at once.
What a Real Voynich Vowel System Must Explain
- Why visible character clustering is unlike a clean plaintext vowel/consonant split.
- Why word edges dominate character behaviour.
- Why q-like openings are so constrained.
- Why minim endings are so common.
- How Currier A/B/C-like regimes affect the mapping.
- Whether labels use the same system.
- Why visible words are short and narrowly distributed.
- How restored source words form stable morphology.
- How source grammar emerges without free vowel insertion.
- How the system predicts unseen tokens.
A vowel hypothesis becomes interesting when it explains these together.
One character that “looks like a vowel” is almost irrelevant by comparison.
What Survives the Vowel Work
- No accepted manuscript-wide Voynich vowel inventory exists.
- Simple two-state clustering does not robustly produce the clean vowel/consonant split familiar from known alphabetic plaintext.
- The famous Reddy/Knight result instead highlighted word-final position strongly.
- That specific result should not be treated as universal because reproduction and initialisation sensitivity remain concerns.
- Word-edge structure is strong enough to confound simple phonological clustering.
- Omitted-vowel, syllabic, abbreviatory and cipher models remain plausible mechanism classes.
- Character segmentation is upstream of vowel classification.
- EVA character names do not imply phonetic values.
- Any vowel system must transfer across independent pages and produce stable source-side phonology and morphology.
- The failure of a clean vowel split is evidence about representation, not proof of meaninglessness.
What Does Not Survive as Established Knowledge
- EVA o is proven to be the vowel /o/.
- EVA a is proven to be /a/.
- The q-o sequence is proven equivalent to familiar Latin-script qu.
- Reddy & Knight proved that Voynich vowels occur only at word ends.
- A later HMM run proved the true Voynich vowel set.
- Voynich is proven to be an abjad.
- Voynich is proven to be a syllabary.
- Minims are proven vowel markers.
- Bench components are proven consonant-vowel combinations.
- The absence of obvious vowel letters proves non-language.
The phonological question survives.
The vowel table does not.
A Better Vowel Analysis
- Define the character representation first.
- Repeat analysis under analytical and synthetic segmentations.
- Model token position explicitly before interpreting hidden states phonologically.
- Test multiple initialisations and report stability.
- Analyse Currier/RZ regimes separately before pooling.
- Separate labels from prose.
- Compare with alphabetic, abjad-like, devowelled, syllabic and abbreviated controls.
- Compare realistic cipher outputs as well as plaintext.
- Require vowel restoration to improve morphology and grammar.
- Freeze the mapping before testing unseen bifolia.
- Keep failures and unstable clusterings visible.
What Would Count as a Real Vowel Breakthrough?
Imagine a character model is developed independently from stroke composition and distribution.
Under that model, a three-state rather than two-state phonological analysis produces stable classes across many random initialisations.
One class behaves vowel-like in Currier A, B and intermediate material.
The same class appears inside labels.
A proper-name crib independently fixes several sound values and confirms the class.
When the mapping is applied to unseen prose, legal syllable patterns and recurring morphemes emerge in one historically plausible language.
Restoring omitted vowels, where required, follows fixed grammatical rules and reduces rather than expands ambiguity.
The resulting source text predicts words on pages whose pictures were not used in the derivation.
That would be a genuine vowel breakthrough.
The decisive moment will not be when one Voynich glyph is called a vowel. It will be when a stable vowel system makes thousands of other glyph relationships easier to predict.
Primary School: Remove the Vowels
Write:
THE CAT SAT ON THE MAT
Remove most vowels:
TH CT ST N TH MT
A familiar reader may still reconstruct the sentence.
An outsider has a much harder task.
The exercise shows how a text can encode vowels implicitly without writing separate vowel symbols every time.
Lower Secondary: One Sign Can Carry a Syllable
Create six invented symbols for ka, ke, ki, ko, ku, kə.
Ask students to find the “vowel letters”.
There are none as independent symbols.
Vowel information is real.
It is simply bundled into larger units.
This demonstrates why failure to identify standalone Voynich vowels does not settle whether the text represents speech.
Upper Secondary: Let the Model Discover Position Instead of Vowels
Create an artificial token system in which one family occurs almost only at word endings.
Run a simple two-state clustering exercise.
The model may separate endings from non-endings even if the underlying language has vowels everywhere.
Students learn that an unsupervised statistical state must be interpreted after its distribution is understood.
A cluster label is not a discovered meaning.
JC and Adult Readers: Identifiability Before Phonology
At a higher level, the vowel problem is an identifiability problem.
Several latent mechanisms can produce similar observable bigram statistics.
- phonological vowel/consonant classes;
- prefix/core/suffix positions;
- syllabic components;
- abbreviation states;
- cipher states.
A two-state model cannot identify which mechanism generated the clustering merely from the existence of two states.
Additional independent observables are required.
Proper-name phonology.
Morphological alternation.
Cross-regime transfer.
Source-language grammar.
The correct question is therefore not “which state is vowel?”
It is “which hidden mechanism makes this state structure inevitable?”
A Parent and Teacher Guide
- Never infer sound from the EVA letter used to name a glyph.
- Separate visible characters from functional characters.
- Remember that vowels can be omitted or bundled into syllables.
- Treat statistical clusters as descriptive until independently interpreted.
- Control word-edge effects before calling a cluster phonological.
- Test multiple Currier/RZ regimes.
- Require restored vowels to create stable source grammar.
- Keep unstable and failed classifications visible.
The general reasoning lesson is:
when a familiar category fails to appear, first ask whether the representation hides it before concluding the underlying phenomenon is absent.
Reader Checklist: Before You Identify a Voynich Vowel
- What counts as one character?
- Are compounds split or merged?
- Has token position been controlled?
- Is the result stable across model initialisations?
- Does the class recur in Currier A and B?
- Does it recur in intermediate/RZ-C material?
- Does it work in labels?
- Is an abjad-like model being compared?
- Is a syllabic model being compared?
- Is abbreviation being compared?
- Are realistic cipher controls included?
- Does the proposed vowel class produce plausible phonotactics?
- Does it create recurring morphology?
- Do fixed values predict unseen text?
- Would the conclusion survive another reasonable transliteration?
Frequently Asked Questions
Has anyone identified the Voynich vowels?
No accepted manuscript-wide vowel inventory exists. Different statistical approaches have produced different tentative classifications.
What did Reddy and Knight find?
Their 2011 two-state HMM produced an unusual separation in which one state strongly tracked word-final characters rather than an ordinary vowel class. Later commentary notes that the result is sensitive and difficult to reproduce robustly.
Does that mean Voynich has no vowels?
No. Vowels may be omitted, encoded inside larger units, represented through abbreviation, transformed by cipher or otherwise obscured by the visible representation.
Could Voynich be an abjad?
It remains a plausible mechanism class. A real abjad-like solution must reconstruct vowels under consistent linguistic rules and produce coherent source language across unseen text.
Could one Voynich glyph represent a syllable?
Yes in principle. No accepted syllabary or consonant-vowel decomposition has been demonstrated.
Is EVA o actually the vowel o?
No. EVA names are transliteration conventions chosen for visual convenience and do not assign sound values.
Why does the vowel question matter?
Because the presence, absence or bundling of vowels affects word length, phonotactics, entropy, morphology and which language/cipher models remain plausible.
What is the strongest current conclusion?
Voynich character statistics do not currently yield a clean, robust plaintext-like vowel/consonant partition. That is a genuine constraint on simple alphabetic readings, but it leaves several structured language and encoding mechanisms open.
Related eduKateSG Reading
Continue Through the Voynich Research Map
Research and Further Reading
- Sravana Reddy & Kevin Knight — What We Know About the Voynich Manuscript (2011)
- René Zandbergen — Voynich Character Analysis and Vowel/Consonant Classification
- Claire Bowern & Luke Lindemann — The Linguistics of the Voynich Manuscript
- Lindemann & Bowern — Character Entropy in Modern and Historical Texts
- René Zandbergen — Voynich Transliteration and Character Units
The Final Idea
The vowel problem is useful because it catches us at the moment when a symbol becomes too familiar.
EVA writes o.
Our mind hears /o/.
EVA writes a.
Our mind hears /a/.
But the manuscript never promised us those sounds.
When statistical models try to discover vowel-like classes without those visual suggestions, the result remains unstable and strangely dominated by position.
That may eventually turn out to be because vowels are missing.
Because they are bundled.
Because the script abbreviates.
Because the text is encoded.
Or because we have not yet found the right functional units.
The important thing is that the manuscript has made the simple alphabet model pay a cost.
The Voynich vowel problem will be solved when one representation makes phonology, morphology, word length and unseen text become simultaneously easier—not when one glyph is finally given a familiar vowel name.