Natural languages do not run on nouns alone.
They need glue.
Articles.
Pronouns.
Prepositions.
Conjunctions.
Auxiliaries.
Particles.
Small words that may carry little concrete meaning by themselves but organise relationships among the larger words.
English has the, of, and, to, is.
Latin has its own grammatical machinery.
Italian another.
Languages with richer morphology may push more grammatical information into endings rather than separate words.
Some languages omit categories that English treats as obligatory.
But every natural language still needs ways to express grammatical relations.
So where is that machinery in Voynich?
There are high-frequency tokens.
There are recurring families.
There are positional constraints.
There are word-class-like statistical clusters.
There are relationships across token boundaries.
Yet no accepted analysis has demonstrated that one Voynich form means and, another means of, another is an article, another an auxiliary, and that the same assignments survive the manuscript.
The Function-Word Problem is not whether Voynich contains frequent words. It is whether any recurring structural class behaves like grammatical glue under rules that remain stable when the page, topic and text regime change.
Quick Read
Direct answer: Voynich contains high-frequency recurring tokens and statistical classes that could include grammatical function words, but no accepted function-word inventory exists. A genuine function-word layer should be relatively stable across semantic topics, recur in syntactically coherent positions, interact predictably with neighbouring token classes, survive Currier and scribal variation under an explainable mechanism, and improve unseen-text prediction without first assigning lexical meanings from pictures.
- Function words are grammatical operators rather than content-heavy lexical items.
- They are often frequent and broadly distributed, but frequency alone does not identify them.
- Voynich has common tokens such as daiin-like and chedy-like forms, but their Currier distributions differ strongly.
- A universal function word should not disappear entirely merely because the illustration changed, unless the grammar or encoding state changed too.
- Some languages encode functions morphologically rather than as separate words, so Voynich grammatical glue may live in prefixes, suffixes or token edges.
- Visible spaces may not map perfectly to lexical words, which complicates function-word discovery.
- Reddy & Knight showed that useful structural classes can be inferred without knowing meanings, making class-level syntax more important than exact-token guesses.
- The newer Edge Problem makes token beginnings and endings especially important candidate locations for grammatical information.
- Currier A/B and newer intermediate classifications must be controlled rather than averaged away.
- Labels should not be expected to use the same amount of grammatical glue as prose.
- Record-like Quire 20 text may use reduced or formulaic grammar.
- A cipher can preserve function-word behaviour, distort it, split it among homophones or encode it inside larger groups.
- A generated system can create frequent structural markers that mimic function words.
- A real function-word hypothesis must predict syntactic environments before semantic translation.
- No accepted Voynich analysis currently meets that standard manuscript-wide.
Function Words Are Usually Small but Structurally Expensive
Function words are easy to overlook because many are short.
Yet remove them from English and the language becomes skeletal.
“Plant grows garden.”
We can guess.
But did the plant grow in the garden?
Was it brought to the garden?
Is it the plant already discussed?
Did one event happen and another?
The little words carry relationships.
A technical manuscript needs relationships too.
Ingredient with operation.
Object with property.
Quantity with substance.
Condition with instruction.
If Voynich encodes language, that relational machinery must exist somewhere.
The First Trap: “Most Frequent” Does Not Mean “Function Word”
In many languages, the most frequent words are grammatical.
That makes frequency a useful starting heuristic.
It is not a definition.
A botanical catalogue may repeat “root”.
A medical manual may repeat “water”.
A codebook may repeatedly use a control marker.
A generator may repeatedly prefer one central template.
Therefore a frequent Voynich token becomes function-word-like only if its distribution reflects grammatical role more strongly than topic or production mechanics.
Currier A and B Make Naive Frequency Guessing Fail Immediately
One of Voynich’s most important constraints is that the most common visible forms change by text regime.
Traditional Currier A favours forms that are much less common in B.
Traditional B favours chedy-like families that can be absent or rare in A.
René Zandbergen’s summaries emphasise that Currier differences occur at characters, character groups and whole words, and that newer RZ classification adds an intermediate C-like state.
This means the highest-frequency token in one regime cannot simply be declared the manuscript’s equivalent of “the”.
If it is grammatical, we need to explain the regime dependence.
- different language;
- different dialect;
- different morphology;
- different cipher state;
- different scribal convention;
- different document role.
The function-word theory inherits the Currier Problem rather than bypassing it.
Function Can Survive Surface Variation
A crucial idea follows.
The same grammatical function does not need one visible form everywhere.
English has am, is, are, was, were.
Different surface forms share a grammatical family.
A cipher may give one plaintext function word several homophones.
A dialect may change its spelling.
A scribe may abbreviate it differently.
Therefore the correct unit may be a function class, not one exact token.
This is why statistical word classes matter.
Reddy & Knight Shift the Question From Words to Classes
Reddy and Knight’s 2011 work is valuable because it asks what can be recovered before decipherment.
Rather than requiring known word meanings, distributional methods can group forms according to where they occur.
Words occupying similar environments may belong to related grammatical or structural classes even when their exact values are unknown.
This suggests a better function-word search.
Do not ask only:
Which Voynich word means “and”?
Ask:
Which recurring forms occupy a narrow set of grammatical-looking contexts across otherwise changing lexical material?
That question can be attacked without a dictionary.
A Function Word Should Travel Better Than a Content Word
If the manuscript contains multiple topics, content vocabulary should change.
Plant names on plant pages.
Astronomical terms on diagrams.
Preparation terms near vessels.
Function words should often travel further because grammar is reused across topics.
This creates a direct test.
Find forms with broad cross-section distribution but constrained local context.
Those are stronger grammatical candidates than words concentrated in one visual province.
Not all languages behave identically, and document genres may reduce grammar, but cross-topic stability remains a useful signal.
Long-Range Memory and Function Words Are Complementary
The Long-Range Memory Problem asks what textual structure persists beyond immediate neighbours.
Function words provide one mechanism by which local grammar can stay stable while lexical content drifts.
If a candidate function class survives across distant pages while content families turn over, that is language-like.
If every high-frequency class drifts with local vocabulary, grammatical interpretation becomes harder.
Maybe Voynich Has Few Independent Function Words
English encourages us to expect many separate grammatical words.
That expectation is not universal.
Highly inflected languages encode relationships inside words.
Case endings can replace prepositions.
Verb endings can encode person and number.
Clitics can attach grammatical items to neighbouring words.
Therefore Voynich’s grammatical glue may be embedded in the very word-family components we have been calling prefixes and suffixes descriptively.
This links the Function-Word Problem to Word Families without duplicating it.
Word Families owns visible internal relatives.
This article asks which of those components behave grammatically across different lexical cores.
The Edge Problem Makes Attached Grammar Even More Plausible
Recent 2026 work reports a striking boundary asymmetry: whole exact token identity predicts little, while glyphs immediately across token boundaries remain coupled.
If replicated, that is exactly where clitics, agreement markers or other edge grammar could hide.
But language is only one explanation.
- cipher state transfer;
- scribal joining convention;
- weak spaces splitting larger units;
- structured generator rules.
The useful point is narrower.
A function-word search that considers only whole tokens may miss grammar living at token edges.
Articles May Be Absent Even in Real Language
English speakers expect “a” and “the”.
Latin does not have an article system equivalent to modern English.
Therefore failure to find article-like tokens tells us little by itself.
A strong analysis must compare grammatical categories appropriate to candidate source languages rather than searching for English one-to-one equivalents.
Prepositions May Become Case Endings
A relation such as “of”, “to” or “from” can be expressed by a separate preposition.
Or by case morphology on a noun.
Or by both.
If Voynich reflects a heavily inflected language, searching for one tiny high-frequency token meaning “of” can fail because the relation is inside the word.
The correct test becomes distributional.
Does one recurring ending appear on tokens occupying similar relational contexts?
That is a grammar question before it is a translation question.
Conjunctions Should Reveal Coordination
A conjunction such as “and” often links similar grammatical objects.
Noun and noun.
Clause and clause.
Ingredient and ingredient.
This gives candidate conjunctions a structural test.
A form proposed as “and” should often sit between similar token classes or parallel structures.
If it appears indiscriminately inside every position, the conjunction hypothesis weakens.
No semantic dictionary is needed to test parallel-class preference.
Auxiliaries Should Prefer Verb-Like Environments
If Voynich has separate auxiliary verbs, they should interact with one latent class more than others.
That class might eventually prove verb-like.
Again, we do not need to know whether a particular token means “is”.
We need to show that it behaves like a grammatical operator over a stable neighbouring class.
Pronouns Should Track Discourse Differently From Names
Pronouns are frequent but referential.
They typically occur in prose, not as isolated labels attached to objects.
This gives labels a useful control.
A candidate pronoun that is equally common as a star or plant label becomes suspicious.
Document role can therefore rule out some grammatical interpretations before translation.
Labels Should Contain Less Grammatical Glue
Labels are usually fragments.
Name.
Identifier.
Quantity.
Coordinate.
They often omit the grammatical machinery of full sentences.
If a proposed function-word family appears frequently in running prose but is depleted in labels, that is positive role evidence.
If it is equally abundant in labels, perhaps it is orthographic or cryptographic instead.
Quire 20 May Use Formulaic Grammar
The star-marked pages look record-like.
Records often compress grammar.
“Rose — warm — morning — two drachms.”
This is meaningful without full prose syntax.
Therefore lack of ordinary function-word frequencies in Quire 20 would not disprove language.
It might identify a telegraphic register.
The function-word analysis must respect document type.
Ciphertext Can Preserve or Hide Function Words
A simple substitution preserves the frequency and distribution of function words strongly.
A homophonic cipher can split one frequent function word into several visible variants.
A nomenclator can replace one function or phrase with a code.
Nulls can obscure boundaries.
Transposition can move grammatical operators away from their source neighbours.
Therefore failure to see familiar high-frequency words does not eliminate ciphertext.
But every cipher layer buys complexity that must later collapse back into stable source grammar.
A Generated System Can Mimic Grammatical Glue
A token generator may use control markers.
Prefix A selects one template.
Marker B starts an entry.
Ending C balances a line.
These repeated structural items can look like function words statistically.
The discriminator is whether they support a recoverable semantic syntax or merely a production syntax.
Grammatical-looking distribution is not yet grammar.
The Best Candidate Function Words Should Be Boring
This sounds odd.
But function words tend to be structurally reusable.
A candidate that occurs only on spectacular zodiac diagrams is more likely content-specific.
A candidate that appears across dull prose, different topics and many pages while keeping similar contextual behaviour is more interesting grammatically.
The grammatical glue should survive when the content gets exciting.
A Function-Word Theory Needs Cross-Page Conservation
Suppose a token family is proposed as an article or conjunction.
Measure its context on one subset of pages.
Freeze the contextual prediction.
Then move to unseen bifolia.
Does the same family still sit between the same latent classes?
Does the distribution survive another scribe?
If Currier changes the surface form, can one transformation recover the same grammatical role?
That is what turns a local statistical curiosity into grammar.
What Survives the Function-Word Work
- Voynich contains frequent recurring tokens and token families.
- Frequency alone does not identify grammatical function.
- Currier A/B strongly changes high-frequency visible forms.
- Function may therefore live at the class level rather than one exact token.
- Natural-language grammar can be expressed by independent words, affixes, clitics or combinations.
- Labels provide a useful negative control because they should often contain less full-sentence grammar.
- Record-like text may compress function-word usage without ceasing to be meaningful.
- The Edge Problem makes token boundaries and affix-like elements especially important.
- Cipher and generated systems can mimic some function-word statistics.
- A real grammatical class must generalise across unseen text and predict neighbouring classes.
- No accepted manuscript-wide Voynich function-word inventory currently exists.
What Does Not Survive as Established Knowledge
- daiin is proven to mean “the”.
- chedy is proven to be a function word.
- The most common Voynich token must be an article.
- Voynich must contain an English-like article system.
- Every grammatical relation must be written as a separate token.
- Statistical word classes are proven parts of speech.
- High-frequency edge glyphs are proven clitics.
- Function-word-like structure proves natural language.
Grammatical structure remains a live question.
The dictionary of grammatical words does not.
A Better Function-Word Analysis
- Separate exact tokens from token families and edge classes.
- Measure cross-topic distribution.
- Measure contextual selectivity, not frequency alone.
- Control Currier/RZ regime.
- Control proposed scribal hand.
- Separate labels, prose and record-like entries.
- Test affix/clitic models as well as independent words.
- Compare with cipher and generator controls.
- Freeze candidate grammatical classes before semantic interpretation.
- Predict contexts on unseen bifolia.
- Require any eventual translation to preserve the same grammatical behaviour.
What Would Count as a Real Function-Word Breakthrough?
Imagine unsupervised analysis identifies three recurring Voynich token families that occur widely across prose topics but rarely as labels.
One family consistently precedes a latent noun-like class.
One links parallel latent classes.
One interacts with a verb-like class.
The distributions are learned without image semantics.
The classes are frozen.
They predict contexts on unseen bifolia.
Currier A/B surfaces differ, but one compact transformation maps both onto the same underlying grammatical roles.
A later decipherment then assigns historically plausible function meanings under one source language and the grammar remains stable.
That would be a genuine breakthrough.
The first convincing Voynich function word will not be the form that looks most like “the”. It will be the class that keeps doing the same grammatical job even when the manuscript changes everything around it.
Primary School: The Invisible Glue
Write:
cat table
Then:
the cat is under the table
Ask which new words tell us the relationship.
Children see why small grammatical words can carry large structural information.
Secondary School: Find Function Without Meaning
Create an invented language in which one symbol always appears between two nouns.
Do not tell students what it means.
They can still infer a grammatical role from distribution.
This models Voynich analysis before translation.
JC and Adult Readers: Closed-Class Discovery
At a higher level, function-word discovery is a latent-class problem under distribution shift.
The candidate class should have:
- high contextual reuse;
- lower topic specificity;
- strong neighbour-class selectivity;
- relative stability across domains;
- limited class inventory.
Voynich adds regime and segmentation uncertainty, so the model must marginalise over or explicitly control those variables.
The goal is not to cluster everything.
It is to identify a small closed class whose structural behaviour transfers.
Reader Checklist: Before You Call a Voynich Form a Function Word
- Is the claim based only on high frequency?
- Does the form travel across topics?
- Does it prefer a stable neighbouring class?
- Is it depleted in labels?
- Does it survive Currier A/B as one function or require unrelated values?
- Could it instead be an affix or clitic?
- Could it be a cipher control marker?
- Could it be a generator template marker?
- Does the contextual rule work on unseen bifolia?
- Does a later semantic reading preserve the previously discovered structural role?
Frequently Asked Questions
Has anyone identified a Voynich word meaning “the” or “and”?
No manuscript-wide assignment has been accepted as a demonstrated decipherment.
Are the most frequent words probably grammatical?
They are worth testing, but topic vocabulary, cipher markers and generated templates can also be frequent. Contextual behaviour matters more than frequency alone.
Could grammar be inside Voynich words rather than separate words?
Yes. Affixes, clitics or structured token edges are plausible locations for grammatical information.
Why are Currier A and B important?
Because common surface forms differ sharply. Any grammatical interpretation must explain whether the underlying function stays constant while the representation changes.
What is the strongest current conclusion?
Voynich contains enough recurring class-level and boundary structure to make grammatical-function analysis meaningful, but no accepted closed class of function words has yet been recovered.
Related eduKateSG Reading
Research and Further Reading
- Sravana Reddy & Kevin Knight — What We Know About the Voynich Manuscript
- Claire Bowern & Luke Lindemann — The Linguistics of the Voynich Manuscript
- René Zandbergen — Currier A and B: Two Different Languages?
- René Zandbergen — Voynich Text Analysis
The Final Idea
Voynich research often reaches for nouns first.
Which plant?
Which star?
Which body part?
Which ingredient?
But language may yield to smaller, less glamorous things first.
The invisible glue.
A class that links.
A marker that selects.
An ending that agrees.
A grammatical operation that keeps working when the topic changes.
If Voynich encodes language, its grammar should eventually become visible not because one mysterious word suddenly translates beautifully, but because the same small structural operators keep doing the same jobs everywhere else.