VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | Everything eduKate Knows and Tested | EVA, Transcription and the Segmentation Problem

Before you can decipher the Voynich Manuscript, you have to make a quieter decision.

What exactly are you looking at?

One glyph?

Two connected glyphs?

A ligature?

A flourish?

A scribal variant?

A damaged stroke?

A meaningful space?

A slightly larger gap caused by the writer lifting the pen?

These questions sound like housekeeping.

They are not.

Every frequency table depends on them.

Every entropy measurement depends on them.

Every word-count depends on them.

Every attempt to compare Voynichese with Latin, Italian, German, Hebrew, Arabic, a cipher or a constructed system depends on them.

Before meaning comes representation.

And representation is not neutral.

When we turn a medieval handwritten mark into a modern character on a computer, we have already made a theory about where one unit ends and another begins.

This is why EVA—the Extensible Voynich Alphabet—is both enormously useful and easy to misunderstand.


Quick Read

One-sentence answer: EVA is a modern transliteration alphabet designed to represent Voynich glyph forms consistently using ordinary alphabetic characters, but EVA letters are shape labels rather than decoded sounds, and every computational result built from them inherits choices about glyph identity, ligatures, variants, spaces, damaged marks and text layout.

  • EVA means Extensible Voynich Alphabet; it was developed by Gabriel Landini and René Zandbergen with important contributions from Jacques Guy.
  • The ordinary letters used in EVA are convenient labels for Voynich shapes, not claims about pronunciation or plaintext identity.
  • EVA o does not mean the Voynich glyph is the Latin letter O.
  • EVA daiin, qokedy or any other familiar-looking string is a transliteration, not a translated word.
  • Voynich research has used several transliteration alphabets and files over time, including FSG, Currier/D’Imperio, EVA, v101, Takahashi and newer alignment systems.
  • Different transcribers can disagree about whether a stroke is one glyph, two glyphs, a ligature, a rare variant or damage.
  • Extended EVA includes rare and unusual forms that basic EVA cannot represent neatly.
  • Capitalisation and bracketing conventions can encode connected glyphs or ligatures in EVA-based files.
  • Takeshi Takahashi produced one of the first essentially complete EVA transliterations, enabling broad computational work.
  • Interlinear files preserve multiple transcribers’ readings side by side instead of silently choosing one.
  • IVTFF, the Intermediate Voynich Transliteration File Format, records not only text but also page and locus metadata.
  • IVTFF can distinguish running paragraph text, labels, circular text and radial text—important because layout role may affect textual behaviour.
  • Visible spaces provide useful token boundaries but are not proven lexical word boundaries.
  • Changing segmentation can change character counts, entropy, bigram frequencies, token families and apparent language likeness.
  • A serious result should state which transliteration, alphabet, segmentation and treatment of uncertain readings produced it.

The lesson is not “transcriptions are unreliable”.

The lesson is better:

a transcription is a measurement instrument, and every instrument has resolution, assumptions and uncertainty.


Why We Need a Transliteration at All

Imagine comparing a thousand handwritten Voynich words manually.

You zoom one page.

Copy a shape.

Search another page visually.

Repeat.

This works for a handful of examples.

It does not scale to manuscript-wide statistics.

A transliteration converts recurring handwritten forms into a symbolic representation that can be:

  • searched;
  • counted;
  • sorted;
  • compared;
  • clustered;
  • modelled computationally;
  • shared between researchers.

This is a profound transformation.

The manuscript image becomes data.

But the transformation is not automatic.

A human—or a carefully designed recognition process—must decide which visible form corresponds to which symbolic unit.

That decision is the segmentation layer.


Transliteration Is Not Translation

This distinction is so important that it deserves repeating even after our earlier writing article.

Translation says:

this expression means “water”.

Transliteration says:

I will represent this visible shape with the modern symbol o so that we can refer to it consistently.

The modern symbol is an address.

It is not the medieval meaning.

This is easy to forget because EVA produces strings that look pronounceable.

qokeedy.

chedy.

daiin.

Our brains want to sound them out.

That pronunciation is an artefact of the Roman letters chosen for convenience.

EVA is not telling us that Voynichese sounded like those strings.


Why EVA Uses Familiar Letters

A transliteration system has practical goals.

It should be typable.

Searchable.

Convertible.

Readable by software.

Compatible with earlier work where possible.

EVA was designed as a superset capable of representing older transliteration traditions and previously omitted material while using ordinary alphabetic characters.

That practical design is one reason it became so widely used.

But convenience creates cognitive contamination.

If a Voynich glyph is labelled o, it starts feeling circular, vowel-like and familiar.

If another is labelled k, it starts feeling consonantal.

These intuitions come from our alphabet, not from the manuscript.

A disciplined reader mentally replaces “EVA letter” with “shape code”.


Basic EVA and Extended EVA Solve Different Resolution Problems

The manuscript contains common forms and rare forms.

A compact alphabet can represent the common inventory efficiently.

Rare, compound and unusual characters require more resolution.

This is why extended systems exist.

Basic EVA is useful for ordinary text analysis.

Extended EVA preserves distinctions that would otherwise be flattened.

The trade-off is familiar from every representation system.

Compress too much and meaningful variation disappears.

Distinguish every tiny variation and the alphabet explodes into thousands of idiosyncratic forms.

The challenge is deciding which differences belong to identity and which belong to handwriting variation.

This is not a Voynich-only problem.

Every OCR system, palaeographic transcription and speech-recognition system faces a version of it.


One Stroke or Two? This Is the Segmentation Problem

Look at connected handwriting.

A loop touches a vertical stroke.

Is that one glyph?

Two glyphs written closely?

A ligature representing a third unit?

A decorative connection?

Different segmentation choices create different data.

If AB is one glyph, the alphabet has one rare character.

If A+B are two glyphs, the corpus has a bigram.

Those alternatives affect:

  • alphabet size;
  • character frequency;
  • conditional entropy;
  • bigram probability;
  • word length;
  • morphological analysis;
  • cipher-model fit.

The statistics do not merely describe the handwriting.

They describe the handwriting after segmentation.

Before asking what a statistic means, ask what representation created the statistic.


Ligatures Make the Problem Visible

EVA-based transliterations include ways to indicate connected characters and ligatures.

Capitalisation or bracketing can mark forms whose strokes connect.

This is not merely typographic nicety.

It preserves evidence that two component-like forms were written as one connected graphical event.

Later analysis can then choose whether that connection is:

  • graphically important only;
  • a meaningful grapheme distinction;
  • a scribal habit;
  • an encoding feature.

Flattening connected and unconnected forms too early destroys that option.

Preserving every connection forever can also make analysis unwieldy.

The best representation keeps enough information to allow later questions rather than forcing all future questions to inherit one early decision.


The Gallows Characters Show Why Shape Families Matter

Several tall Voynich forms are conventionally called gallows characters.

The nickname describes appearance.

Some also occur in compound or pedestal-like configurations.

A transcription must decide whether these forms belong to:

  • one family with positional variants;
  • separate characters;
  • combinations of simpler strokes;
  • abbreviation-like compounds.

The answer changes statistical interpretation.

For example, paragraph-initial gallows behaviour can look like a special character distribution.

If some variants are actually compounds, part of the effect may belong to graphical composition.

This does not make positional effects unreal.

It tells us the causal explanation depends on how the writing system is segmented.


Rare Characters Are Where Over-Segmentation Becomes Dangerous

Voynich contains numerous extremely rare forms.

Some occur once.

What are they?

  • real rare characters;
  • ordinary characters written unusually;
  • corrections;
  • damage;
  • pen slips;
  • ligatures;
  • later marks.

If every unusual stroke becomes a new symbol, the alphabet becomes artificially large.

If every unusual stroke is collapsed into a common symbol, genuine rare distinctions disappear.

This is a classification problem under uncertainty.

A good transliteration system should preserve uncertainty rather than pretend the rare form was obvious.

This is why extended character inventories and uncertainty conventions matter.

The rarest data points are often where a representation system reveals its philosophy.


Spaces Are Visible—but “Words” Are Still a Hypothesis

The manuscript visibly separates many strings with spaces.

These space-delimited units are enormously useful.

We can count them.

Compare them.

Study their frequencies.

Find local vocabulary.

But calling them words adds a linguistic assumption.

A visible space might separate:

  • lexical words;
  • syllabic groups;
  • morphemes;
  • cipher groups;
  • generated chunks;
  • scribal rhythm units.

Therefore “token” or “word-like unit” is often the safer analytical term.

This does not reduce the usefulness of spaces.

It keeps their job bounded:

spaces give us an observable segmentation convention; they do not yet tell us which linguistic unit the convention represents.


A Space Error Can Manufacture a Word Family

Imagine a string ABCDEF.

One transcriber sees:

ABC DEF

Another sees:

AB CDEF

Now suppose ABC appears often elsewhere.

Under the first transcription, it becomes a common token.

Under the second, it disappears.

One tiny space decision changes vocabulary statistics.

The same is true for prefixes and suffixes.

If spacing is inconsistent, an apparent morphological family may partly reflect scribal spacing practice.

This is why serious analysis benefits from checking uncertain spaces against images and against multiple transcribers.

Word statistics are not wrong.

They are conditional on a word-boundary representation.


Takeshi Takahashi Made the Corpus Computationally Usable at Scale

One of the major practical advances in Voynich research was the production of an essentially complete EVA transliteration by Takeshi Takahashi.

Completeness matters.

Early studies could focus on selected sections or partial transcriptions.

A manuscript-wide file allows researchers to ask:

  • Where does a token occur across the entire book?
  • How do Currier regimes differ?
  • How local is a vocabulary item?
  • What happens at paragraph openings?
  • How do labels differ from running text?
  • Which rare characters concentrate on particular folios?

This is one reason transliteration infrastructure deserves to be treated as scholarship rather than clerical work.

A research field can only ask questions at the resolution its data representation supports.

Better corpus infrastructure expands the question space.


Multiple Transcriptions Are a Feature, Not an Embarrassment

Why not choose the best transcription and discard the rest?

Because disagreement contains information.

If five transcribers agree on one glyph, confidence is high.

If they disagree, the image location deserves inspection.

An interlinear file can preserve these readings side by side.

This lets analysis distinguish:

  • robust features that survive transcription choice;
  • fragile features driven by one transcriber’s segmentation;
  • systematic differences between transcription philosophies.

A result that appears in every major transcription is stronger than one that disappears when ambiguous glyphs are represented differently.

Disagreement is therefore not merely noise to average away.

It is an uncertainty map.


The Interlinear File Is a Way of Keeping Reality Attached to the Data

Gabriel Landini’s collection of major older transliterations into an interlinear format made an important conceptual move.

Instead of pretending there was one unquestionable text file, multiple readings could remain aligned by location.

This preserves two things at once:

  • a computational representation;
  • the fact that the representation came from human readings of a physical image.

The alignment matters because two transcribers may disagree at exactly one glyph while agreeing on the surrounding line.

The disputed locus can then be revisited rather than turning into a corpus-wide silent inconsistency.

This is good data engineering and good epistemology for the same reason:

uncertainty remains addressable.


IVTFF Adds the Missing Question: Where Is This Text?

A plain text file can tell us character order.

Voynich requires more.

Is the text:

  • inside a paragraph;
  • an isolated label;
  • written around a circle;
  • written along a radius;
  • inside a star-marked entry;
  • on a particular folio with a particular Currier classification?

The Intermediate Voynich Transliteration File Format, IVTFF, preserves location and metadata alongside transliteration.

Its locus types can distinguish running paragraphs, labels, circular text and radial text.

This is crucial because the same glyph sequence may behave differently by document role.

A label is not a tiny paragraph.

Circular text may have spatial order constraints different from horizontal prose.

Metadata restores dimensions that a flat character stream would otherwise destroy.


Labels Need Their Own Textual Population

An isolated string beside a zodiac figure may be a name.

A string inside a paragraph may be a verb.

If we pool them without distinction, frequency analysis can mix textual roles.

This is why locus metadata matters.

We can compare:

  • label vocabulary versus paragraph vocabulary;
  • zodiac labels versus plant-fragment labels;
  • diagram labels versus starred-entry text.

If one token concentrates in labels, that may suggest a naming or coding role.

If it occurs everywhere, perhaps it is grammatical or structurally common.

The key point is not the final semantic answer.

It is that layout creates a conditional distribution worth preserving.

Flattening layout loses evidence.


Circular Text Creates an Ordering Problem

Horizontal prose has an obvious local direction.

Circular text creates new questions.

Where does the sequence begin?

Clockwise or anticlockwise?

Does the circle have a start at all?

Is one ring independent from another?

A transliteration has to linearise the circle somehow.

That linearisation can create an artificial first token.

If a study analyses “sentence beginnings” in circular text without accounting for this, the result may partly come from editorial convention.

This is a beautiful example of representation changing topology.

A circle must be cut somewhere to become a line.

The cut is ours unless the manuscript marks it.


Radial Text Creates Another Geometry

Some diagram text runs along radii.

Now position can mean:

  • distance from centre;
  • direction;
  • association with a sector;
  • relationship to a ring;
  • ordinary writing order.

A plain corpus can preserve only character sequence unless metadata retains the geometry.

This is why the best Voynich representation is not simply text transcription.

It is text-plus-location.

Eventually, a semantic theory may need text-plus-visual-object-plus-physical-sheet as well.

The more of reality we preserve, the fewer future questions are foreclosed by the representation.


Character Counts Depend on the Alphabet

How many distinct Voynich characters are there?

The answer depends on the representation.

A coarse alphabet merges variants.

A fine alphabet separates them.

A stroke-level alphabet may decompose what another alphabet treats as one glyph.

Therefore “Voynich has N characters” is not purely an observation.

It is:

under segmentation scheme S, we identify N character classes.

This wording is less dramatic and more accurate.

The distinction matters especially when comparing alphabet size to natural languages, syllabaries or cipher systems.

A comparison made with incompatible unit definitions can create false similarity or false difference.


Entropy Depends on Representation

Voynich character-level entropy has attracted enormous attention because the script is highly predictable in unusual ways.

But entropy is computed over symbols.

Change the symbols and the entropy changes.

Merge two character classes.

Split one compound.

Treat ligatures differently.

Resolve uncertain glyphs.

The measured sequence changes.

This does not invalidate entropy research.

It tells us what robust entropy research should do.

  • state the alphabet;
  • state the corpus;
  • state segmentation decisions;
  • test sensitivity to alternative reasonable representations where possible.

A phenomenon that survives multiple representations is much stronger evidence about the manuscript than one that exists only under one transcription convention.


Word Length Depends on Spaces and Glyph Segmentation Together

Voynich word-like units are famously constrained in length and structure.

But word length is not a raw visual fact.

It depends on:

  • where spaces are placed;
  • how many glyphs a compound contains;
  • whether rare connected forms are split;
  • how uncertain characters are handled.

A five-glyph token under one alphabet can become a four-glyph or six-glyph token under another.

Therefore comparisons with word-length distributions in Latin or Italian should be made at compatible representational levels.

If Voynich visible glyphs encode syllables while Latin comparison uses letters, raw length comparison becomes semantically uneven.

The same visual string can sit at a different linguistic depth.

Representation defines the ruler.


Currier A/B Is Strong Partly Because It Survives Representation Changes

The Currier A/B distinction is not based on one fragile glyph.

It appears across multiple distributional features and has been rediscovered under later computational analyses.

This makes it a useful example of robustness.

Exact cluster boundaries can vary with feature set and transliteration.

But the broad non-uniformity of the manuscript remains.

That is what we want from a serious Voynich phenomenon:

not independence from representation—nothing measured is—but stability across reasonable representations.

Robustness is one way the world pushes back against our encoding choices.


A Proposed Decipherment Must Say Which Text It Is Deciphering

This sounds almost absurd.

Yet decipherment claims often begin from a transcription file without discussing its assumptions.

A serious method should specify:

  • which transliteration alphabet;
  • which transcription file;
  • how uncertain glyphs are handled;
  • how ligatures are segmented;
  • how spaces are treated;
  • whether labels and running text are pooled;
  • how circular text is linearised.

Otherwise another researcher cannot reproduce the input.

And if the input is not reproducible, the output cannot become strong evidence.

A translation rule that depends on one transcriber’s uncertain character should say so.

Confidence should propagate from input to conclusion.


Image-Based Work and Transcription-Based Work Need Each Other

There are two bad extremes.

Extreme one: ignore transcription and work only by visual intuition.

This does not scale.

Extreme two: treat the transcription file as if it were the manuscript itself.

This forgets the physical strokes and uncertainty that produced the data.

The stronger relationship is cyclical.

image → transliteration → analysis → suspicious result → return to image.

If a rare statistical feature depends on six unusual glyphs, inspect those six glyphs.

If transcribers disagree in exactly the region driving a theory, lower confidence.

If the feature survives image review, confidence increases.

The representation should remain correctable by the object.


New Alignment Alphabets Show the Representation Problem Is Still Active

Voynich transcription did not end with EVA.

More recent work has developed supersets and analytical alignment alphabets designed to compare major transliteration traditions more systematically.

The STA approach, for example, was designed to capture distinctions across Extended EVA and v101.

An analytical alignment alphabet goes farther by organising strokes and minims into families.

Why build another alphabet after decades of research?

Because the underlying representation question is not fully closed.

Different alphabets preserve different kinds of visual identity.

The continuing work is not evidence of failure.

It is evidence that the field is learning to preserve uncertainty and interoperability more carefully.


The Representation Layer Can Create False Discoveries

Suppose a researcher discovers that character X almost always follows character Y.

Exciting.

Then image inspection shows X and Y are usually two components of one connected glyph that the transcription split.

The “grammar rule” was partly created by segmentation.

Or suppose a token appears only in Currier B.

Then another transcription merges one glyph variant and the token becomes a common form found across A and B.

The semantic distinction weakens.

These are hypothetical illustrations of a general danger:

some patterns belong to the manuscript; some belong to the way we encoded the manuscript.

The job of robust analysis is to tell them apart.


The Representation Layer Can Also Hide Real Discoveries

Compression has the opposite danger.

If two visually distinct glyph variants are merged, a real positional system may disappear.

If labels and paragraph text are pooled, a genuine label vocabulary may vanish into general frequency.

If circular text is linearised without geometry, sector relationships disappear.

If rare characters are normalised away, a special notation system may become invisible.

So the ideal representation has two competing goals:

  • simplify enough to analyse;
  • preserve enough to remain faithful.

That balance is the central art of transcription.

Too little abstraction and computation becomes impossible.

Too much abstraction and the manuscript disappears inside our model.


A Good Representation Must Be Reversible Enough to Return to the World

This is perhaps the deepest principle in the entire transcription problem.

If a computational result says:

this rare form occurs 17 times

we should be able to return to those 17 physical locations.

Look at them.

Ask whether the transcriptions are visually comparable.

See whether damage or layout explains some cases.

A representation that loses the route back to the manuscript becomes dangerous because errors cannot be corrected by the source object.

This is why page IDs, loci and interlinear alignment are so important.

Data should remain answerable to parchment.


What EVA and Transliteration Can Tell Us

  • Which represented glyph classes recur.
  • How often represented sequences occur.
  • Where represented tokens concentrate.
  • How line and paragraph positions differ.
  • How Currier regimes differ statistically.
  • How labels differ from running text when metadata is preserved.
  • Where transcribers agree or disagree.
  • Which patterns survive alternative representations.
  • Which parts of the manuscript deserve return-to-image inspection.

That is enough to support serious structural linguistics without a single translated sentence.


What EVA Cannot Tell Us Alone

  • How the manuscript sounded.
  • Whether a glyph represents a letter, syllable, morpheme or cipher group.
  • Whether spaces are lexical words.
  • What any EVA string means.
  • Whether a ligature is semantically distinct.
  • Which language underlies the text.
  • Whether the text is plaintext or ciphertext.
  • Whether visual variants are functional or scribal.
  • The exact original segmentation intended by the writer.

EVA is an interface to the problem.

It is not the answer encoded in Roman letters.


A Better Representation Audit

1. Name the source images

Which scan set and folio locations underlie the analysis?

2. Name the transliteration

Takahashi, Zandbergen-Landini, reference file, another corpus?

3. Name the alphabet

Basic EVA, extended EVA, v101, STA or another mapping?

4. State segmentation rules

Ligatures, compounds, spaces, uncertain glyphs and rare characters.

5. Preserve layout classes

Paragraphs, labels, circular text and radial text should not be pooled automatically.

6. Test robustness

Does the main result survive another reasonable transliteration or segmentation?

7. Inspect high-leverage disagreements

Return to images where transcription choice drives the conclusion.

8. Carry uncertainty forward

A doubtful input should not produce a certain semantic conclusion.

This is enough to make public research more reliable without exposing any private research machinery.


What Would Count as a Representation Breakthrough?

Imagine high-resolution palaeographic work establishes that several forms currently treated as separate glyphs are systematically composed from smaller stable units.

A new segmentation dramatically simplifies the character inventory.

The new representation explains previously unusual entropy.

It predicts positional constraints.

Independent transcribers can apply it reliably.

That would be a major breakthrough even before translation.

Or imagine the opposite: supposedly connected components prove to be distinct graphemes with stable independent distributions.

Again, the model of the writing system changes.

A representation breakthrough tells us what the units are.

That may be exactly what a later decipherment needs before it can tell us what the units mean.


Primary School: A Label Is Not the Thing

Draw a strange symbol on paper.

Tell a child:

We will call this symbol “B”.

Then ask:

Does that mean the symbol makes the sound “b”?

No.

“B” is only the classroom label.

That is EVA in miniature.

The child learns that names we assign for convenience do not automatically reveal the identity of the thing named.


Lower Secondary: One Shape or Two?

Draw two connected loops.

Ask one group of students to treat them as one symbol.

Ask another group to treat them as two.

Give both groups a page of repeated strings and ask them to count frequencies.

Their results differ.

Nothing in the page changed.

The representation changed.

This is one of the most important lessons in data science:

the categories used to measure reality can change the pattern we measure.


Upper Secondary: Run the Same Statistic Under Two Segmentations

Give students a toy corpus.

In Version A, connected forms are one symbol.

In Version B, they are two symbols.

Calculate:

  • alphabet size;
  • most common bigram;
  • average token length;
  • character predictability.

The numbers move.

Now ask which conclusions survive both versions.

Those surviving conclusions are more robust.

The exercise teaches sensitivity analysis without needing advanced mathematics.


JC and Adult Readers: Representation Is Part of the Model

At a higher level, the Voynich segmentation problem is a reminder that data are not raw reality.

Reality is parchment, ink, strokes, gaps and layout.

A transliteration extracts discrete symbolic units.

A statistical analysis operates on those units.

A linguistic model interprets the statistics.

Every layer can introduce assumptions.

The mature question is therefore not merely:

What does the model say?

It is:

Which transformation from physical manuscript to model input made this statement possible?

That question is useful far beyond Voynich.

It belongs anywhere measurements are produced from representations.


A Parent and Teacher Guide

EVA and transcription are an excellent way to teach students what happens before data analysis.

  1. Separate label from identity. EVA letters name glyph shapes for analysis; they do not pronounce them.
  2. Ask what counts as one unit. One stroke, one glyph, one ligature or two glyphs?
  3. Preserve uncertainty. If transcribers disagree, the disagreement is data.
  4. Keep layout. Labels, circular writing and paragraphs may behave differently.
  5. Check sensitivity. Does the conclusion survive another reasonable representation?
  6. Return to the source. High-leverage statistical claims should be traceable back to the manuscript image.
  7. Carry uncertainty forward. A doubtful glyph cannot support a certain translation.

This is how students learn that data cleaning is not a boring stage before reasoning.

It is already reasoning.


Reader Checklist: Before You Trust a Voynich Statistic or Translation

  1. Which transliteration alphabet is used?
  2. Which transcription file?
  3. Are EVA letters being treated as sound values?
  4. How are ligatures and connected glyphs represented?
  5. How are rare characters handled?
  6. How are uncertain readings marked?
  7. Are spaces assumed to be word boundaries?
  8. Are labels pooled with paragraph text?
  9. How is circular or radial text linearised?
  10. Does the result survive another major transcription?
  11. Can the claimed examples be traced back to specific folio loci?
  12. Does one ambiguous segmentation drive the conclusion?
  13. What part of the finding remains if the segmentation changes?

If a paper or decipherment does not answer these questions, the result may still be interesting.

Its evidentiary resolution is simply lower than its prose may suggest.


Frequently Asked Questions

What is EVA?

EVA is the Extensible Voynich Alphabet, a modern system for transliterating Voynich glyph forms with typable alphabetic codes.

Does EVA decode the manuscript?

No. EVA labels visible forms. Its Roman letters are not known phonetic or semantic values.

Who developed EVA?

It was developed by Gabriel Landini and René Zandbergen, with important contributions and suggestions from Jacques Guy, as part of a broader effort to make Voynich transliteration complete, interoperable and computationally useful.

Who is Takeshi Takahashi?

He produced one of the first essentially complete EVA-based transliterations of the manuscript, which became widely used in computational research.

Why are there multiple Voynich transcriptions?

Handwritten glyphs can be ambiguous, and different systems make different segmentation choices. Multiple transcriptions let researchers identify which results are robust and where the manuscript is genuinely difficult to represent.

What is IVTFF?

It is the Intermediate Voynich Transliteration File Format, which stores transliterated text together with page and locus metadata and supports several transliteration alphabets.

Are spaces real words?

The spaces are real visual features and give useful token boundaries. Whether those boundaries correspond exactly to lexical words is unresolved.

Why does segmentation matter so much?

Because changing what counts as one glyph or one word changes the sequence used for frequencies, entropy, word lengths, morphological families and cipher comparisons.

What is the safest way to use EVA?

Treat it as a reversible analytical representation tied to specific manuscript locations, preserve uncertain readings, keep text roles separate and test important findings across alternative reasonable transcriptions.


Related eduKateSG Reading


Research and Further Reading


The Final Idea

It is tempting to think the mystery begins after transcription.

First we type the manuscript.

Then the real work starts.

The Voynich Manuscript teaches the opposite.

The real work has already started when we decide what one glyph is.

It continues when we decide what one space means.

When we flatten a circle into a line.

When we decide a connected form is one symbol or two.

When we call a label a word.

Every one of those decisions creates a representation.

And the representation becomes the world our statistics can see.

That is why EVA is so valuable.

Not because it solved Voynich.

Because it made the manuscript measurable while still leaving a route back to the parchment.

The best representation is not the one that makes the mystery disappear. It is the one that lets reality correct us when our representation is wrong.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading