VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | The Phonotactics Problem: Could Voynichese Actually Be Pronounced?

If Voynichese is language, could a human mouth actually say it?

That question sounds almost embarrassingly basic.

Yet it sits underneath nearly every proposed decipherment.

A script can look strange and still encode ordinary speech.

A cipher can destroy visible pronunciation while preserving recoverable plaintext underneath.

A shorthand can compress syllables.

An abjad can omit many vowels.

A syllabary can package consonant and vowel information into one sign.

So the visible page does not have to be directly pronounceable.

But any claim that Voynich ultimately encodes natural language eventually owes us a sound system or an explicit reason why sound cannot be reconstructed directly from the visible symbols.

The Phonotactics Problem asks whether one stable representation of Voynich can generate legal, repeatable sound sequences rather than merely produce words that resemble a target language after enough adjustments.


Quick Read

Direct answer: no accepted pronunciation system for Voynichese exists. The visible script is unusually constrained within tokens, no robust manuscript-wide vowel/consonant split has been recovered, and different character segmentations produce different apparent sound structures. Those facts make naive letter-by-letter pronunciation weak, but they remain compatible with omitted vowels, syllabic writing, abbreviation, lost phonemic distinctions, cipher transformation or another non-transparent orthography.

  • Phonotactics concerns which sound sequences a language permits.
  • Languages constrain syllables even when spelling systems look very different.
  • A proposed Voynich sound system should produce stable onset, nucleus and coda patterns or another historically plausible syllable architecture.
  • The Vowel Problem comes first because vowel information may be omitted, merged or bundled.
  • Bowern and Lindemann argue that Voynich’s low conditional character entropy may reflect lost phonemic distinctions among other possibilities.
  • That means one visible glyph may represent a broader class than one ordinary phoneme.
  • Strong within-token positional restrictions can resemble phonotactics, morphology, abbreviation or code grammar.
  • q-like openings and minim-heavy endings are highly constrained but not securely phonetic.
  • Currier A/B differences may reflect pronunciation, orthography, morphology, cipher state or topic.
  • Labels provide a useful test because proper names often preserve phonological structure strongly if they are actually names.
  • A syllabic or shorthand model predicts different unit boundaries from an alphabetic model.
  • A cipher model may hide phonotactics completely at the visible layer while restoring it after decryption.
  • A generated system may imitate syllable-like restrictions without speech.
  • A real pronunciation should make unseen tokens more—not less—predictable under one fixed phonological grammar.

Phonotactics Is Not the Same as Pronunciation Guessing

Suppose we assign EVA o the sound /o/ because it looks like o.

Assign EVA a /a/.

Assign a gallows /k/.

Now read a token aloud.

That is pronunciation guessing.

Phonotactic analysis asks something stricter.

Do the resulting sequences behave like a coherent language?

  • Which sounds can begin a word?
  • Which can end it?
  • What clusters are legal?
  • Where must a vowel occur?
  • How many syllables can one token contain?
  • Do the same restrictions recur across thousands of tokens?

A real sound system constrains more than one attractive word.

Syllables Give Us a Useful Neutral Scaffold

Many spoken languages organise words into syllables.

A simple model is:

onset + nucleus + coda

The onset often contains consonants.

The nucleus often contains a vowel.

The coda may contain consonants.

Languages differ enormously.

Some allow complex clusters.

Others prefer simple CV syllables.

Some permit syllabic consonants.

The scaffold matters because Voynich tokens themselves appear highly zoned.

Optional beginnings.

Restricted cores.

Characteristic endings.

That can resemble syllable architecture.

It can also resemble morphology or code slots.

The Vowel Problem Is Upstream

Before we can identify syllables, we need to know where vowel information lives.

The dedicated Vowel Problem article showed why this is unresolved.

Simple statistical clustering does not produce a stable plaintext-like vowel/consonant partition.

That leaves several possibilities.

  • vowels are omitted;
  • vowels are encoded in compound glyphs;
  • vowels are represented by context rather than separate signs;
  • the units are syllables rather than letters;
  • the visible script is ciphertext;
  • the system is not directly phonographic.

Phonotactic analysis must therefore run under several representation models rather than assume the answer.

Low Character Entropy May Mean We Lost Distinctions

Lindemann and Bowern compared Voynich character predictability with hundreds of modern and historical language samples.

The manuscript remains unusually predictable under substantial transcription and glyph-composition changes.

One interpretation they discuss is loss of phonemic distinctions.

Imagine a writing system that does not distinguish several sounds that speech distinguishes.

The visible sequence becomes more constrained because different spoken possibilities collapse into one written class.

This gives the phonotactics problem a deep twist.

Voynich may fail to look like ordinary spelling not because it lacks language, but because its orthography may preserve fewer phonological distinctions than the spoken system underneath.

That remains a hypothesis class, not a recovered pronunciation.

An Abjad-Like Model Changes the Visible Syllable

Suppose many vowels are not written.

A visible sequence such as KTB can correspond to several vocalised words depending on language and grammar.

Voynich tokens would then look consonant-heavy.

Visible word lengths would shrink.

A naive vowel detector would fail.

But a real abjad-like solution has to do something difficult:

restore vowels under rules constrained by morphology and syntax rather than freely.

If every consonant skeleton can become whatever target word is needed, phonotactics has not been recovered.

A Syllabary Changes the Character Count

Suppose one visible unit represents an entire syllable.

Then a five-glyph Voynich token could represent five syllables rather than five letters.

Or compound forms could encode consonant-vowel pairs.

The correct phonotactic model would then operate over syllable classes, not individual EVA characters.

This is exactly why segmentation matters.

Count a bench as one unit and the syllable theory changes.

Split it into components and another theory becomes possible.

Shorthand Creates a Third Representation

A shorthand system may encode syllables, consonant skeletons, abbreviations and special whole-word forms together.

That means visible glyph-to-sound mapping can be mixed.

One sign may be syllabic.

Another an abbreviation.

Another a positional modifier.

This is why the Shorthand Problem deserves a separate owner.

Phonotactics then becomes a downstream test: does the expanded shorthand recover a coherent historical sound system?

Ciphertext Can Hide Phonotactics Almost Completely

Homophony can split one plaintext sound among several ciphertext signs.

Verbose substitution can expand a sound into a group.

Nomenclators can replace whole words.

Transposition can move phonological neighbours apart.

Therefore visible Voynich phonotactics may be weak even if source-language phonotactics is ordinary.

A cipher theory does not escape the sound problem.

It postpones it.

After decryption, a coherent source sound system should emerge.

q Is a Perfect Phonotactic Trap

EVA q almost always occurs in a strongly constrained opening pattern.

Readers trained on Latin script immediately think of qu.

That resemblance may be meaningless.

EVA names are visual labels.

If q is phonological, its near-obligatory partner and strong token-initial preference should fit a source-language phonotactic rule.

If it is grammatical, the same distribution has another cause.

If it is cipher state, another.

The observation is strong.

The sound value is not.

Minim Endings Could Be Codas—or Not

Repeated-stroke strings such as EVA iin-like endings cluster toward token ends.

That looks coda-like.

It could reflect:

  • phonological endings;
  • grammatical suffixes;
  • abbreviations;
  • numerical or quantitative notation;
  • cipher groups;
  • scribal terminal forms.

A coda interpretation should make one sound-system prediction across many roots, not merely note that endings are endings.

Word Families Can Be Phonological Families

Natural-language words related by morphology often remain phonologically related.

One suffix changes the final syllable.

A prefix adds an onset.

Voynich’s dense families could therefore conceal real phonological paradigms.

But generation and cipher models can create the same visual families.

The key discriminator is whether a proposed sound change has a stable grammatical or lexical effect across many unrelated bases.

Currier A and B Could Encode Different Phonotactics

If A and B are dialects or languages, phonotactic differences are expected.

If they are cipher states, visible phonotactics may differ while source phonotactics remains constant.

If they are scribal conventions, the differences may be orthographic only.

This gives any pronunciation model a powerful test.

Does one underlying syllable inventory explain A, B and intermediate material through small transformations?

Or does each regime require a new sound system?

The latter is possible but increasingly expensive.

Labels Should Be a Late-Stage Phonological Test

If labels are names, they may preserve pronunciation strongly.

If a stable Voynich phonology is discovered in prose, apply it to labels without modification.

Do repeated syllable patterns emerge?

Do historical proper-name forms become plausible?

If labels require a completely separate pronunciation system, either document role is different or the model is weak.

Labels should validate a mature phonology, not be mined endlessly for names to invent one.

Pronounceable Is Not the Same as Correct

Humans are extremely good at making unknown strings sound pronounceable.

Assign enough vowels.

Merge awkward clusters.

Call one sign silent.

Use historical spelling variants.

Soon nearly any token can be spoken.

That proves little.

A correct phonology should constrain pronunciations sharply enough that many alternatives become impossible.

The Sound System Must Be Economical

Suppose one glyph can represent six consonants.

Another can be any vowel.

A third is silent when necessary.

A word can be rearranged.

Now thousands of readings become possible.

The sound system is not discovered.

It is underconstrained.

Every ambiguous sound assignment should count as model complexity.

A Real Phonology Should Predict New Words

Learn the phonotactic grammar on one subset.

Freeze it.

Now inspect unseen Voynich tokens.

The model should tell us which pronunciations are legal.

Which are impossible.

Which syllable boundaries are likely.

If every new word forces new phonemes or exception rules, the system is not converging.

What Survives the Phonotactics Work

  • No accepted Voynich pronunciation or phoneme inventory exists.
  • Voynich token-internal character placement is unusually constrained.
  • A simple alphabetic vowel/consonant model has not yielded a clean stable partition.
  • Loss of phonemic distinctions is one plausible explanation for unusually low character entropy.
  • Abjad-like, syllabic, shorthand and cipher models all remain viable representation classes.
  • Positional q-like and minim-like behaviour can be phonological but is not proven to be.
  • Currier differences are a major test for any proposed sound system.
  • Labels should eventually validate a stable phonology if they contain names or ordinary linguistic units.
  • Pronounceability alone is cheap; predictive phonotactics is strong evidence.

What Does Not Survive as Established Knowledge

  • EVA letters have their Latin-alphabet sounds.
  • EVA q is proven /k/ or /kw/.
  • Minim endings are proven codas.
  • Voynich is proven consonant-only writing.
  • Voynich is proven syllabic.
  • One pronunciation that yields plausible words is evidence of decipherment.
  • Currier A/B are proven dialect phonologies.
  • Low entropy proves loss of phonemes.

A Better Phonotactic Analysis

  1. Define the functional character representation first.
  2. Test alphabetic, abjad-like, syllabic and abbreviation-aware models.
  3. Infer syllable boundaries without using desired translations.
  4. Measure legal onset/nucleus/coda distributions or equivalent structure.
  5. Control Currier/RZ regime and scribal hand.
  6. Test labels separately.
  7. Count ambiguous sound values as complexity.
  8. Freeze the phonological grammar.
  9. Predict pronunciations of unseen tokens.
  10. Require any later lexical decipherment to obey the previously discovered sound system.

What Would Count as a Real Phonotactic Breakthrough?

Imagine a segmentation model identifies a compact set of functional units.

Without using a candidate language dictionary, those units fall into stable onset-like, nucleus-like and coda-like classes.

The same classes recur in A and B with one small systematic transformation.

Held-out tokens receive legal syllabifications automatically.

A later independent crib fixes several sound values and lands exactly inside the predicted classes.

Then source-language morphology and proper names begin to emerge without adding new phonemes freely.

That would be a phonotactic breakthrough.

Voynich becomes pronounceable scientifically when one sound grammar tells us in advance which readings are impossible—not when we discover that enough imagination can make every token speak.

Primary School: Legal and Illegal Sounds

Ask a child which sounds more like an English word:

  • blan;
  • nglb.

Neither needs to be a real word.

Yet one obeys familiar sound rules better.

This is phonotactics: knowing what could be a word before knowing what it means.

Secondary School: Hide the Vowels

Write a sentence without most vowels.

Students can often restore several possible pronunciations.

Then give grammatical context.

The possibilities shrink.

This shows why omitted-vowel writing can remain readable to insiders but difficult to decipher externally.

JC and Adult Readers: Phonology as a Latent Constraint System

At a higher level, a proposed phonological interpretation should maximise explanatory compression.

A small inventory of latent sound classes should account for large numbers of visible token sequences.

The model should outperform nonphonological slot grammars on held-out prediction rather than simply fit observed tokens.

The hard question is identifiability: the same visible positional structure can arise from speech, morphology, shorthand, cipher or generation.

Independent evidence must break that equivalence.

Reader Checklist: Before You Pronounce Voynich

  1. What counts as one functional character?
  2. Where is vowel information?
  3. Is the model alphabetic, abjad-like, syllabic or mixed?
  4. Are silent signs allowed?
  5. How many sound values can each unit have?
  6. What syllable structures are legal?
  7. Does the model work across Currier A/B?
  8. Does it work on labels?
  9. Does it predict unseen token syllabification?
  10. Do candidate words obey the phonology without one-off exceptions?

Frequently Asked Questions

Do we know how Voynichese was pronounced?

No accepted pronunciation system exists.

Could Voynich omit vowels?

Yes in principle. Omitted-vowel models remain viable but must restore vowels consistently under source-language grammar.

Could each Voynich glyph be a syllable?

Possibly, but no accepted syllabary has been demonstrated. Compound-glyph and shorthand models complicate the unit question further.

Does low entropy mean Voynich lost phonemic distinctions?

It is one plausible interpretation discussed in linguistic work, not an established identification of the writing system.

What is the strongest current conclusion?

Voynich has strong token-internal structure but no accepted mapping from that structure to a stable sound system. Any linguistic solution must eventually explain that gap.

Related eduKateSG Reading

Research and Further Reading

The Final Idea

Voynich has always invited voices.

Readers silently assign sounds to its shapes almost as soon as they see them.

That instinct is human.

Science asks for something harder.

A sound system that survives the manuscript.

One that explains why openings are openings.

Why endings are endings.

Where vowels went.

Why A and B differ.

Why labels behave differently.

And why thousands of unseen tokens still obey the same constraints.

The manuscript will begin to have a voice when one phonological model reduces the freedom of every future reading instead of giving us more ways to make the next word sound familiar.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading