VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Tangential Voynich | Token Boundaries Are Our Cuts

Thesis: Before we can count, compare, translate or model the Voynich Manuscript, we must decide what the units are. Those units are not guaranteed to be supplied by the manuscript itself. A token boundary can be our cut.

This article extends eduKate’s Cognitive Art work on tokens into the Voynich programme. The central warning is simple: the mathematics only sees the units the observer gave it.

Reality Has Fewer Boxes Than Our Minds

A river is easy to name. Its beginning is not always easy to locate. A cloud is easy to see. Its exact edge is not. A species is useful to classify. Evolution does not stop at the line in our taxonomy. A city boundary can be legally precise while the lived city extends far beyond it.

Human reasoning works by imposing useful cuts on continuous or complicated reality.

CONTINUOUS / COMPLEX REALITY
→ CUT
→ TOKEN
→ MANIPULABLE MODEL

The cut is often necessary. It is not therefore natural.

A Token Is a Decision About What Counts as One Thing

Tokenisation asks: where does one unit end and the next begin?

In ordinary English, spaces make the answer feel obvious. Yet even modern language contains contractions, compounds, hyphenation, punctuation, emojis, numbers and multi-word expressions that complicate the boundary.

In an unidentified script, the uncertainty is much larger.

  • Is a visible stroke one glyph or part of a ligature?
  • Is a repeated pair two glyphs or one functional unit?
  • Is a blank space a lexical boundary?
  • Are line breaks linguistic, operational or merely graphic?
  • Are labels and running text generated by the same rules?
  • Does one apparent token contain smaller recurrent units?

Each answer creates a different mathematical universe.

The Statistics Inherit the Cut

Suppose we decide that every blank-delimited string is a word.

Now we can calculate:

  • word frequency;
  • word length;
  • vocabulary size;
  • type-token ratios;
  • Zipf-like frequency structure;
  • word-to-word transitions;
  • topic models;
  • network relationships.

Every result looks objective.

But all of it is conditional on the first cut.

SPACE = WORD BOUNDARY
↓
WORD INVENTORY
↓
STATISTICAL MODEL
↓
INTERPRETATION

If the first equality is wrong, the later mathematics can be exact while the conceptual object is wrong.

This is why the EVA segmentation problem is foundational rather than technical housekeeping.

The Same Trap Exists in the Pictures

Tokenisation is not only textual.

Look at a plant-like page. Where is the object boundary?

Is the entire drawing one object? Are the root, stem, leaves and flower separable components? Are the colours meaningful classes? Is a band around a stem part of the plant or a diagrammatic operator? Are several components drawn together because they belong biologically, procedurally or compositionally?

The moment we call the whole image “a plant,” we have performed an object tokenisation.

The current visual-grammar lane therefore does something deliberately less satisfying. It separates branch events, terminal forms, nodes, bands, lobes, channels and attachment geometries before asking what the compound object is.

Components before objects is token hygiene.

Sections Are Tokens Too

“Herbal section.” “Astronomical section.” “Biological section.” “Pharmaceutical section.” “Recipes.”

These are useful catalogue handles.

They are also large-scale tokens imposed on the manuscript.

The frozen eduKate matrix already shows why caution is needed. Physical bifolios, proposed hands, Currier regimes, visual morphologies and distributional clusters overlap without collapsing into a single segmentation.

That means one page can belong simultaneously to several different structural systems.

PAGE
∈ PHYSICAL STRUCTURE
∈ SCRIBAL STRUCTURE
∈ VISUAL STRUCTURE
∈ DISTRIBUTIONAL STRUCTURE

A one-dimensional “section” token can therefore erase multidimensional membership.

The Boundary Creates the Question

Once we decide where a unit ends, we decide which questions become available.

If the unit is a word, we ask about vocabulary.

If the unit is a glyph, we ask about character entropy.

If the unit is a line, we ask about line-initial and line-final effects.

If the unit is a page, we ask about page clustering.

If the unit is a bifolio, we ask about physical coherence.

If the unit is a visual component, we ask about recurrence across apparently different objects.

The unit does not merely structure the answer. It structures the research programme.

Retokenisation as an Experimental Method

Tangential Voynich therefore treats retokenisation as a deliberate test.

Take the same data and represent it at multiple scales:

  • stroke;
  • glyph;
  • multi-glyph unit;
  • blank-delimited token;
  • line;
  • paragraph;
  • page;
  • bifolio;
  • quire;
  • visual component;
  • visual assembly.

Then ask which measured relationships survive scale changes.

A pattern that exists only under one fragile tokenisation should be treated differently from a pattern that reappears under several plausible schemes.

Boundary Sensitivity

This can be formalised as sensitivity analysis.

Let T be a tokenisation rule and M be a measured property.

M = M(T)

Now vary T within a plausible family.

If M changes dramatically, the result is representation-sensitive.

If M remains stable, the result has stronger claim to being an invariant of the underlying object.

This does not tell us what the pattern means. It tells us how much the pattern depends on our cut.

Why Humans Forget Their Own Cuts

Useful tokenisations become cognitively sticky.

After years of reading, “word” feels like a thing rather than a convention supported by a writing system. After years of botany, “leaf” feels like an obvious unit even when developmental biology reveals continuous structures. After years of maps, national borders look inevitable.

This is a normal feature of expertise. Compression makes thought efficient.

It also creates Normalcy Blindness.

The expert stops seeing the tokenisation because the tokenisation has become the background of thought.

Tangential Systems Force Different Cuts

This is one reason far-away systems are useful.

A linguist may cut Voynich into words. A composer may cut it into motifs. A compiler engineer may cut it into operators, operands and scopes. A railway engineer may cut it into nodes and transitions. An ecologist may cut it into populations and niches.

Most of those cuts will be wrong as historical descriptions.

But comparing them exposes which structures are robust to the cut.

A wrong tokenisation can be useful if it teaches us which observations do not depend on it.

The Token Audit

Before any Tangential claim is promoted, ask:

  1. What exactly is the unit?
  2. Who defined its boundary?
  3. Is the boundary visible, inferred or inherited?
  4. Would another reasonable analyst cut it differently?
  5. Does the effect survive an alternative cut?
  6. Does the unit name smuggle function into the observation?
  7. Can the unit be described neutrally?
  8. What evidence would demonstrate that the boundary is functional rather than merely graphic?

World Return

The Voynich Manuscript may contain letters, words, plants and sections.

It may also contain units for which those names are poor approximations.

The first obligation is therefore not to abolish familiar units.

It is to stop confusing their familiarity with proof.

Every time we draw a boundary around Voynich, we should ask whether we discovered a unit or created one.

Tangential Voynich keeps changing the cut because a structure that survives our cuts is more interesting than a structure produced by them.


Continue: The Inverse Representation Problem · The Representation Trap · EVA, Transcription and the Segmentation Problem

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading