VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | edkSG Vol. 6 — The Held-Out Confounding Check

Voynich | edkSG Research Volumes · EDKSG-VOY-V006 · Edition 1.0.0 · 5 September 2026, Singapore

Research status: executed held-out-folio exploratory test on a documented ZL3b-derived transcription representation. Not a full-corpus replication, independent palaeographic review, translation or decipherment.

Previous: Vol. 5 — The Matched-Context Check · Research record · Voynich Research Library

The strongest result in this volume is not that a Voynich clue became stronger.

It became weaker.

Volume 5 found a clear pooled association in our development excerpts: when the preceding eligible transcription span ended in y, the next eligible form was more likely to begin qo rather than o. But the same article also showed that most of the evidence came from one page and that like-for-like comparisons were scarce.

We therefore moved to new folios that were not part of the earlier development excerpts and asked the same narrow question.

Across five held-out folios—f100r, f100v, f101r, f101v and f102r1—the unadjusted association remains positive. The descriptive pooled odds ratio is approximately 2.69.

Then we compare more like with like: the same folio and the same current form after removing the optional initial q. Seven informative strata containing 29 observations remain. Under a Mantel–Haenszel stratified comparison, the common odds-ratio estimate falls to approximately 1.63; the standard test of a common odds ratio of one gives p ≈ 0.50.

Within this held-out sample, the pooled association survives, but the evidence for an additional predecessor-y effect beyond folio and current-form composition does not.

This does not prove that the relationship is absent. It means the stronger version of the clue has not yet earned promotion.

1. Why this is the correct kind of progress

Voynich research is unusually vulnerable to attractive regularities. The manuscript is structured enough that one can find strong-looking differences almost anywhere: by page, line, token family, scribal assignment, Currier population, illustration class or local position.

The important question is not whether a pattern exists. It is whether the pattern distinguishes its nearest alternatives.

In Volume 5, the obvious alternative was composition. Perhaps preceding y genuinely changes the choice between oR and qoR. Or perhaps pages and current-form families that naturally contain more y endings also naturally contain more qo forms. A pooled count cannot decide between those explanations.

The correct response was not to write a richer interpretation around the 5.03 development-sample odds ratio. It was to seek observations that had not participated in building that result and to preserve the same-form comparison.

That is the purpose of this volume.

2. A new source route closes one access gap but not every one

The original Voynich transcription source remains René Zandbergen’s ZL3b file, identified in its header as version 3b of 13 May 2025. Direct byte acquisition from the original host had repeatedly failed in our earlier environment.

For this test we used a public research repository that preserves a ZL3b source snapshot and deterministic per-folio derived files. The repository documents the raw source as ZL3b-n.txt, records its provenance, and explicitly warns that the transliteration alphabet and glyph grouping remain editorial choices rather than final ground truth. It also documents permission from René Zandbergen for use of the transliteration files with attribution. [2–3]

The repository’s current raw ZL3b object is identified by Git blob 2a4533ab9bdfa85db9bad602d590978953055df1. That Git object identity is useful for reproducibility inside the repository. It is not the earlier SHA-256 value recorded in our frozen workbook, so it does not by itself close the original checksum-comparison job.

For the five held-out folios we used the repository’s diplomatic files, which retain a raw-IVTFF column alongside a cleaned first-alternative column. Their Git object identities are individually recorded by the repository. [4–8]

This source route is therefore stronger than manually copying a few browser lines without an explicit preserved repository object. It is still one represented transcription lineage, not independent observation of the manuscript’s glyphs.

3. We keep the same narrow comparison

The target remains deliberately limited.

For each selected P-type line in the five held-out folios, we examine adjacent spans separated by a confident dot in the diplomatic cleaned representation. A candidate enters the contrast only if the preceding span and current span contain ordinary lowercase alphabetic transcription characters under this simplified exploratory rule.

The current span must be either:

  • oR: begins with o followed by a nonempty remainder, or
  • qoR: begins with qo followed by a nonempty remainder.

The matching key removes the optional q from the latter so that oR and qoR can be grouped as the same current-form family for the contextual comparison.

The predictor is whether the preceding span ends in y. This is not the separate question of whether the current span itself carries final y. It is not a semantic definition of y. It is not a claim that q is linguistically a prefix.

The analysis therefore tests a relationship between represented strings, not a translation rule.

4. The pooled held-out association remains positive

The five folios produce 157 eligible transitions under this exploratory rule.

Preceding spanCurrent qoRCurrent oRTotal
Ends in y242246
Does not end in y3279111
Total56101157

The descriptive pooled odds ratio is:

(24 × 79) ÷ (22 × 32) = 2.693…

This is a weaker association than the 5.03 observed in the earlier development excerpts, but it points in the same direction at the pooled level.

That is useful. A result that disappears completely the moment new pages are introduced would have been an immediate warning that the original pattern was extremely local.

But this is not the decisive check.

The pooled table still allows folio composition and current-form composition to do much of the work.

5. The like-for-like comparison changes the story

We next grouped transitions by folio and by the same current matching key. A group becomes informative for this question only if it contains both current outcomes—oR and qoR—and both predecessor classes—ending in y and not ending in y.

Only seven such groups remain, containing 29 observations.

FolioCurrent matching keyy→qoy→onon-y→qonon-y→on
f100rokeol01203
f100vokeol10225
f101rokeol11158
f101vodaiin01102
f101vokeey11125
f102r1oteol20114
f102r1okol10012

A Mantel–Haenszel comparison across these strata gives a common odds-ratio estimate of approximately 1.63. A standard test of the null value one gives p ≈ 0.498. A homogeneity test across the small strata gives p ≈ 0.126.

These p-values should not be treated as a magical pass/fail machine. The sample was not a random draw from an abstract population, the strata are small, and the representation is one transcription lineage.

The important evidential change is simpler:

The strong pooled association does not remain strong after the comparison is restricted to the same folio and same current-form family.

The data currently support a composition-sensitive association more clearly than they support a general predecessor-y rule.

6. What this does to the earlier hypothesis

We can now separate several statements that should not be merged.

Statement A: y-ending predecessors and qo forms are associated in pooled selected data.
Supported in both the development sample and this held-out five-folio sample.

Statement B: the association is deterministic.
Already defeated by explicit counterexamples in Volume 5.

Statement C: the predecessor ending adds substantial information after controlling for the exact current-form family and folio.
Not established by this held-out test. The stratified estimate is much smaller and the present data do not distinguish it from no additional effect.

Statement D: the pooled association is entirely an artefact of composition.
Also not proved. The current stratified sample is small; failure to establish an additional effect is not proof that the true effect is exactly zero.

Statement E: y or q has a particular semantic meaning.
Not tested. Nothing in this result supplies a translation.

This is a useful narrowing. The research pathway has moved from “there appears to be a rule” toward a more precise question: which part of the apparent rule belongs to the surrounding distribution of forms, and which part remains after those distributions are controlled?

7. Held-out does not mean perfectly independent

The five folios were not part of the earlier 191-record development excerpts used in Volume 5. In that operational sense, they are held out from that analysis.

They are not independent manuscripts. They belong to the same historical object and the same ZL3b transcription lineage. The per-folio files are derived from a common preserved source snapshot. Their cleaning rules share one importer. Their current-form families can recur because the manuscript itself is structured.

Nor did we choose these five folios through a preregistered random sampling scheme. They were available as individually documented diplomatic files in the external benchmark repository and were selected after the previous research question already existed.

So this is a stronger challenge than reusing f115r again. It is not a final validation experiment.

The correct language is held-out exploratory check, not independent replication.

8. The representation itself remains part of the experiment

The external repository is unusually helpful because it documents its transformation choices. Its importer chooses the first listed alternative reading for canonical output, removes IVTFF markup, and makes explicit separator decisions. Its README warns that these are practical benchmark choices rather than a final glyph ontology. [2–3]

For this volume we worked from the diplomatic files and applied a deliberately conservative eligibility rule to their cleaned column. Any span containing nonalphabetic uncertainty residue was excluded from the simplified contrast rather than repaired toward a preferred result.

That does not mean the representation is neutral. A different treatment of doubtful spaces, alternative readings, connected glyphs or character grouping could alter which forms enter the test.

The next stage should therefore not rely on one representation alone if the effect becomes important. A structural claim gains strength when it survives plausible representation changes whose differences were defined before the result was inspected.

This is especially important here because the hypothesis concerns boundaries. Boundary hypotheses are unusually sensitive to the decisions that define a boundary.

9. What the wider research now tells us to demand

The broader Voynich programme has repeatedly converged on one methodological requirement: a clue becomes valuable when it predicts something that its nearest lower-level explanation cannot already predict.

For the present problem, a pooled y→qo association is a lower rung. It demonstrates organisation in the selected representation. It does not yet tell us whether the predecessor ending is the operative variable.

The next rung is incremental prediction. On a genuinely withheld population, compare a context-only model with the same model plus the predecessor-y feature. Use the same eligible observations for both models. Report which groups are excluded because the comparison has no overlap.

If the added feature improves prediction across several physical and production contexts, the structural claim strengthens.

If it does not, the research should stop treating the pooled association as a candidate decoding rule and instead model it as a consequence of composition or another correlated variable.

This is not a loss. A false rule removed is one less corridor to walk repeatedly.

10. The next research job

The strongest next job has three parts.

First: source reconciliation. Compare the repository-preserved raw ZL3b snapshot against the exact source identity recorded in our frozen workbook. A Git object identity is not the same checksum scheme as the archived SHA-256, so the actual bytes must be materialised or otherwise compared in a controlled environment before claiming identity.

Second: corpus-wide overlap mapping. Before fitting another model, identify where the manuscript contains genuine same-form contrasts: the same oR/qoR family occurring under both predecessor-ending classes across multiple folios, hands, Currier populations and physical groups. The denominator is the opportunity to compare, not merely the number of tokens.

Third: frozen evaluation. Separate the data used to choose the matching rule from data used to judge the rule. The present development and held-out exploratory samples have now both been inspected. They belong in the training history, not in a future untouched test set.

The desired future result is not “an odds ratio larger than two”. It is a reproducible improvement in discrimination after the relevant context is already known.

11. Research state after Volume 6

Research itemState after this volume
Pooled predecessor-y / qo associationObserved in development and held-out selected data
Deterministic y→qo ruleContradicted by retained counterexamples
Same-folio/same-current-form held-out comparisonExecuted on 7 informative strata / 29 observations
Stratified common OR≈1.63; present data do not distinguish it from 1
Complete source-byte identity with frozen workbookStill open
Full-corpus incremental predictive testNot completed
Current-final-y experimentStill separate and open
Blind visual-component experimentStill separate and open
Semantic interpretation / deciphermentNo new result

This is the kind of checkpoint the catalogue is intended to preserve. The new volume does not make every number larger. It makes one inference smaller and more defensible.

Sources and reproducibility record

[1] Exact predecessor. Voynich | edkSG Vol. 5 — The Matched-Context Check, EDKSG-VOY-V005.

[2] External source repository. cipher_benchmark Voynich source README. It documents the preserved ZL3b source, derivative files, transformation choices, provenance and representation limits.

[3] ZL3b provenance notes. Voynich ZL3b Source Notes, including source URL, version 3b / 13 May 2025, permission record and import choices.

[4–8] Held-out diplomatic records. f100r · f100v · f101r · f101v · f102r1.

[9] Executed local record. EDKSG-VOY-V006-CONFOUND-01. The retained package contains the 157 admitted event records, seven informative strata, result JSON and calculation summary. The numerical findings in this article refer to that run.

Editorial boundary: Wintour House publishes the bounded result; it does not strengthen the scientific claim. No independent scientific review or decipherment is asserted by this edition.

Collection: Voynich Research Library.

EDKSG-VOY-V006 · Edition 1.0.0 · Published 5 September 2026 · Baseline EDKSG-VOY-B000. Earlier volumes remain unchanged. Later corrections should identify the affected result, representation and edition.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading