Voynich | edkSG Research Volumes · EDKSG-VOY-V008 · Edition 1.0.0 · 5 September 2026, Singapore
Research status: completed retrospective comparison-opportunity map on retained event indices, with a page-level exposure audit across the archived 227-unit frame. Not a complete-transcription analysis, a new validation experiment or a decipherment.
Thirteen page units. Three hundred and seventy-nine eligible transitions. At first, that sounds like a steadily expanding test.
But the more demanding comparison has a different shape. Holding page, recorded hand and current-form family constant leaves 13 groups containing 46 observations with variation in both the predecessor ending and the current prefix outcome. Those observations occupy six page units and only three recorded bifolio groups. [4]
The research is accumulating rows faster than it is accumulating distinct contexts in which the competing explanations can be separated.
A second finding concerns the next sample. The thirteen page units already used in these experiments occupy six groups in the inherited bifolio map. Those groups contain 32 page units altogether. Nineteen additional units therefore share a recorded group with previously analysed material. A new page name would not automatically make them a separate group for evaluation. [3, 4]
This volume maps the opportunity to test a claim, rather than producing another headline odds ratio. It identifies where direct comparisons exist, where they are missing, and where a proposed fresh test could quietly reuse related material.
1. More observations are not always more opportunities to compare
The continuing question concerns represented strings of the form oR and qoR. We ask whether the preceding span’s final y provides information about the next eligible form beyond the recorded context. R is a nonempty residual transcription string, not an identified root or a decoded word.
A pooled table can include every eligible transition. A same-form comparison needs something more particular: occurrences of the same current-form family in comparable contexts, with variation in both the predecessor-ending class and the current outcome.
Imagine repeatedly observing one form after y, always as qoR. That repetition may describe a stable local tendency. It does not, by itself, show what happens to the same family after a non-y predecessor. An additional hundred repetitions of the first situation would not supply the missing contrast.
Now suppose the family also occurs as oR after a non-y predecessor. A direct contrast becomes possible within the specified group. It is still not proof of a causal relationship, because other differences may remain, but the relevant variation is finally present.
Our map records this distinction without deciding the outcome of a future model. The word informative in this article means that a group contains both predecessor classes and both prefix outcomes. It does not mean that all four cells are nonzero, that the observations are independent, or that the group carries semantic information.
Groups without this variation are not worthless. Models that share information across groups can use them under additional assumptions. They simply do not provide the direct within-group contrast defined here.
2. What was combined—and what was deliberately not combined
The first input is Volume 5’s preserving-method event record. The second is Volume 7’s corrected preserving-method event index for five additional page units. The extraction reader, geometry layer and boundary analyser retained in the two packages have matching SHA-256 identities. Their recorded methods are therefore compatible at the checked code level. [1, 2, 4]
Under the primary split policy, Volume 5 contributes 201 events and Volume 7 contributes 178. Their page sets do not overlap. The combined index has 379 distinct admitted event identities.
We do not add Volume 6’s 157 cleaned-column events. Those describe the same five page units later reanalysed in Volume 7, under a different representation policy. Counting both as additional independent observations would confuse repeated analysis with additional source material.
The combined preserved input covers thirteen page units. Twelve contribute an admitted boundary event. The remaining unit, f67v1, was preserved in the geometry work but contributes no event to this marked-paragraph comparison. A zero contribution is recorded rather than allowing the page to disappear from the exposure history.
This run works from the retained, source-addressed analytical indices. Current-form families from Volume 5 are converted to the same SHA-256 identification scheme used by Volume 7. These identifiers support exact equality comparisons without republishing a complete transcription. They are not decoded identities or guarantees that a represented form has one historical meaning.
No new glyph observations, image readings or source transcription records were acquired for the map. We did not rerun the complete upstream extraction. The new work is the crosswalk, comparison inventory, dependence summary and exposure audit over identified retained inputs.
3. The 277-group inventory
Grouping the 379 events by page, recorded hand and current-form family yields 277 groups. We then ask two binary questions of each group: does the predecessor class vary, and does the current prefix outcome vary? [4]
| Variation within the defined group | Groups |
|---|---|
| Neither predecessor class nor prefix outcome varies | 245 |
| Only predecessor class varies | 11 |
| Only prefix outcome varies | 8 |
| Both vary | 13 |
| Total | 277 |
The thirteen groups in the final category contain 46 observations. Those are the observations supporting this particular direct contextual contrast. The other 333 admitted events still belong to the analytical collection; their absence from that contrast must not be misreported as absence of structure or meaning.
Some groups have only one occurrence. Others repeat a situation without varying the relevant feature. Either way, a larger pooled sample does not automatically fill the missing cells of a within-group question.
This gives the next acquisition a more useful target than “find more text”. We need to know where the same families occur under contrasting conditions, whether those contrasts extend across more than one production context, and what the representation can actually tell us about those conditions.
It also limits retrospective interpretation. Because the data have already been inspected, this map is not a preregistered discovery. It describes where a future test needs support; it does not turn the current support into untouched evaluation evidence.
4. The comparisons are concentrated in three recorded groups
The archived matrix assigns a bifolio-group identifier to each page unit. We joined the event index to that field without revising it. A bifolio is a folded sheet forming two leaves, but the particular group assignments here remain inherited analytical descriptions, not physical relationships newly verified by this study. [3, 4]
| Recorded bifolio group | All admitted events | Events in informative same-form groups |
|---|---|---|
| Q1-B1 | 32 | 0 |
| Q1-B2 | 8 | 0 |
| Q9-B1 | 13 | 0 |
| Q19-B1 | 33 | 6 |
| Q19-B2 | 145 | 25 |
| Q20-B2 | 148 | 15 |
| Total | 379 | 46 |
The 25 informative observations in Q19-B2 are distributed across f100r, f100v, f101r and f101v. The six in Q19-B1 come from f102r1. The fifteen in Q20-B2 come from f115r. Thus six page units do not translate into six different recorded bifolio groups.
This is a concentration finding, not a calculation of the effective statistical sample size. We have not established that the three groups are independent of one another, or that every pair inside a group has a particular dependence strength. Shared transcription ancestry, production conditions and recurring forms may cross those boundaries too.
The zeros are equally bounded. Q1-B1, Q1-B2 and Q9-B1 lack the required variation within their selected same-page, same-hand, same-family groups. The calculation does not show that the complete corresponding manuscript sections lack such variation. It describes the retained excerpt coverage.
That distinction turns the result into an acquisition question rather than a negative claim about the book: are the missing contrasts absent from the larger source, or simply absent from what we have admitted so far?
5. Matching position makes the shortage sharper
Adding the same within-location ordinal to the matching key leaves four informative groups containing eight observations. These are not eight independent experiments. They are the remaining observations under a more restrictive grouping of the same known sample. [4]
The ordinal is a position in an analytical span sequence. It is not a measured horizontal coordinate, and it can change when the treatment of a doubtful gap changes. The source-location boundaries are retained, so this comparison does not join separate lines or establish a historical reading path.
This result should not be used to declare position-matching mandatory for every question. It shows the cost of asking this particular stricter question with the present data. The gain in contextual specificity comes with a loss of direct comparison support.
A model may choose a less granular description, share information across positions, or impose a regularising assumption. Those choices can be legitimate. They need to be declared as modelling decisions, not described as observations that the archive already contains.
The useful lesson is that “control for everything” is not a completed method. A proposed control needs an accompanying account of which comparisons remain possible after it is applied.
6. The three spacing policies retain the same concentration
The preserving method has three recorded treatments of doubtful seams: split them into spans; join them; or exclude a candidate pair that touches one. We map all three separately. We do not add their totals together, because they are alternative views of overlapping source material. [1, 2, 4]
| Policy | Admitted events | Informative groups / observations | Recorded bifolio groups supporting those observations |
|---|---|---|---|
| Split doubtful seams | 379 | 13 / 46 | 3 |
| Join doubtful seams | 380 | 13 / 44 | 3 |
| Exclude pairs touching doubtful seams | 335 | 11 / 39 | 3 |
The comparison opportunity remains concentrated in the same three recorded groups under all three treatments. This is robustness of a coverage description, not three independent confirmations of a linguistic mechanism.
The totals change because the policies change which spans and neighbouring pairs are represented. A policy with more events is not automatically better, and a policy with fewer events is not automatically more truthful. Each answers a question defined by its treatment of uncertainty.
We deliberately produce no new combined odds ratio or significance test. The purpose of this checkpoint is to establish the population and its usable contrasts before another model is selected or another effect estimate becomes the centre of attention.
That choice also avoids turning a previously examined sample into an apparently fresh experiment simply by calculating one more statistic.
7. The page-level exposure audit reaches beyond the excerpts
The second output uses the entire archived frame of 227 analytical page units in 52 recorded bifolio groups. This is broader coverage than the event sample, but it is coverage of page metadata, not a new corpus-wide token analysis. [3, 4]
We first mark the thirteen page units present in the retained experiments. We then identify every other unit sharing one of their six recorded groups. That adds nineteen. The resulting set contains 32 page units.
| Status in this audit | Page units | What the status permits us to say |
|---|---|---|
| Present in the retained prior analyses | 13 | Already part of this development history |
| Additional units sharing a recorded group | 19 | Need group-aware review before evaluation use |
| Exposure not established by this audit | 195 | Not certified unseen or reserved as a test set |
| Archived frame | 227 | Analytical page units, not a new physical page count |
Under a proposed rule that keeps complete recorded bifolio groups out of a future evaluation whenever part of the group has been used for development, all 32 units would be held aside. The remaining 195 occupy 46 other recorded groups.
Those 195 are not an approved untouched test set. The broader research programme predates these volumes, and its complete page-by-page exposure history has not been reconstructed here. The metadata frame itself has already been examined. A page absent from the retained excerpt list could still have appeared in an earlier analysis, image review or model-development decision.
The map therefore records uncertainty instead of converting an incomplete memory into a claim of blindness.
8. A different page name can still point into the same group
The inherited matrix places f1r and f8r in Q1-B1. A proposed split using f1r for development and f8r for evaluation has different page identifiers, but it does not have disjoint recorded bifolio groups. Our new split check flags that condition. [3, 4]
Likewise, the thirteen directly used units include parts of Q9-B1. Other foldout panels in that same group do not become a separate production context merely because their panel suffixes differ.
The audit’s nineteen additional related units are f7r, f7v, f8r, f8v, f67r2, f67v2, f68r2, f68r3, f68v1, f68v2, f68v3, f99r, f99v, f102r2, f102v1, f102v2, f104r, f104v and f115v. This is a list derived from the archived grouping, not a new reconstruction of the binding.
The example does not establish that earlier results were inflated by this particular split. We did not run such an evaluation or measure its bias. We have identified a condition that a future group-separated test should detect before it is reported.
Official cross-validation guidance makes the general distinction between random row splitting and group-aware evaluation when observations have group structure. It also warns that preprocessing and model choices must not be learned from the final test material. Our manuscript grouping is a task-specific proposal informed by that principle, not a procedure prescribed by the software documentation. [5]
Even a perfectly group-separated split cannot remove every shared source of dependence. It is one explicit boundary, not an independence certificate.
9. The new check refuses to certify what it cannot know
The executable split check accepts proposed development and evaluation page lists, the inherited frame and the known exposure list. It checks for unknown page identifiers, duplicate entries, shared pages, shared recorded groups and evaluation pages already present in the retained analyses.
A proposal using f1r and f8r on opposite sides returns a hold because the group overlaps. A proposal placing f115r in evaluation returns a hold because that page has already been analysed in this pathway.
A structurally separated example using f1r and f3r does not receive a declaration that the test is independent or unseen. It receives only eligible for exposure review. The broader history still needs checking. These examples exercise the software; they are not selected experimental splits. [4]
This distinction matters because a validation tool can create its own false confidence. If its strongest success label is “verified held-out”, users may assume it checked historical facts that were never supplied to it.
The check instead describes exactly what it established. Passing the fields it can inspect is a reason to proceed to the next review, not a reason to erase the fields it cannot inspect.
10. How this changes the next acquisition
The most useful next evidence is not necessarily the longest page or the page that supplies the largest pooled effect. It is evidence that supplies a missing comparison under a stable source and measurement procedure.
For this boundary question, that means finding where the same current-form families occur under both predecessor-ending classes and with both prefix outcomes. It also means learning how widely those opportunities extend across the manuscript rather than repeatedly returning to the contexts already carrying the result.
A discovery stage can inspect source coverage to design such a test. That inspection must be recorded as development exposure. The selected material cannot then be called untouched solely because its role has changed from discovery to evaluation.
A future prediction study should fix the source edition, representation policy, model contrast, grouping scheme and evaluation population before examining the decisive scores. It should compare a context-only model with the same model plus the predecessor-ending feature on the same eligible observations.
If the target is instead a descriptive corpus census, the rules are different: all available material may be counted, but the resulting account should not be presented as a prospective validation. A census and a held-out prediction study have different purposes.
In either case, the current map gives us an explicit shortage to address. It does not prescribe a desired result or require the predecessor-y hypothesis to survive.
11. What was verified in this run
The audit binds its normalized inputs to SHA-256 hashes and records the parent packages and matching analysis-code identities. It rejects mixed spacing-policy input and refuses an attempt to introduce the unreconciled Volume 6 event population as a third source.
The primary map contains 379 unique event identities. Every event resolves to a page in the archived frame. The 277 group sizes sum to 379, and the four group-variation categories sum to 277. The full exposure partition sums to 227 page units. These accounting checks prevent a count from changing units unnoticed.
Twenty-six software tests passed. They include duplicate and missing identity checks, invalid boolean labels, mismatched page and locus addresses, unknown group and hand handling, policy separation, exposure-state counts and the refusal to label an unassessed page as unseen. [4]
A separate relational implementation used SQLite grouping on the frozen primary event index. It recovered the same thirteen informative groups, 46 observations and three supporting recorded bifolio groups without calling the map’s grouping function. This is an internal implementation cross-check. Both paths use the same inputs and were prepared in this investigation; no independent reviewer is claimed.
The work is retrospective. Preliminary counts were inspected before the final auditable run. The design record is therefore not a preregistration, and no part of this exercise should be described as an untouched scientific validation.
12. What remains open
Direct attempts to obtain the complete origin transcription and a public mirror failed in the working environment. Searching the accessible archive did not supply a verified full-file replacement. This volume therefore does not close the complete-source identity check.
The distinction between the two maps must remain visible: the exposure audit covers the 227-unit metadata frame, while the token-comparison inventory covers only the retained thirteen-page-unit sample. A complete corpus-wide token opportunity map has not been produced.
Nor have the inherited bifolio assignments been newly checked against the object, the three supported groups shown to be independent, or the unassessed pages demonstrated to be unfamiliar to every prior analyst.
No new odds ratio, causal result, language identification, translation or q/y meaning is claimed. The current-token final-y experiment and blind visual-component test remain separate open tasks. Nothing about the new exposure map completes them.
The immediate next gate is exact full-source acquisition and reconciliation. Once that succeeds, the unchanged method can generate a broader opportunity census. A genuinely evaluative split then needs its own exposure review and frozen protocol rather than inheriting a success label from this map.
The useful denominator is the comparison we can actually make
Volume 8 leaves the manuscript’s meaning unresolved. It changes the account of our evidence.
We can now distinguish thirteen represented page units from twelve event-contributing units, 379 admitted transitions from 46 observations supporting the direct contextual contrast, and six supporting page units from three recorded bifolio groups.
We can also distinguish thirteen directly used units from nineteen additional related units—and distinguish the remaining 195 unassessed units from a genuinely untouched test set.
These are not reasons to stop. They are instructions for what the next useful evidence must supply.
The next advance should increase our ability to distinguish competing explanations, not merely increase the number of rows in the archive.
Sources and reproducibility record
[1] Corrected-method predecessor. Vol. 7 — The Representation-Consistency Check, EDKSG-VOY-V007, WordPress post 154116. Its saved preserving-event indices supply the five-page extension. Its correction to the interpretation of Volume 6 remains in force.
[2] Earlier compatible event set. Vol. 5 — The Matched-Context Check, EDKSG-VOY-V005, WordPress post 154103. Its source-addressed records supply the earlier eight-page-unit development sample. No independent replication is inferred from combining the two packages.
[3] Metadata frame. Vol. 2 — The Input Audit and its retained matrix-rows.json, derived from the frozen Voynich_Test_1b_Frozen_Segmentation_Matrix.xlsx. Bifolio, page and classification fields remain inherited assignments. The normalized frame used here retains its parent file identity.
[4] Executed local audit. EDKSG-VOY-V008-OPPORTUNITY-01. The research package contains normalized event indices, the 227-unit metadata frame, exposure states, all group ledgers, policy-specific summaries, executable map and split checks, 26-test log and relational cross-check. These calculations operate on retained analytical representations; full-source extraction is not rerun.
[5] Evaluation-method reference. scikit-learn: Cross-validation—evaluating estimator performance, especially preprocessing and grouped-data evaluation, consulted 5 September 2026. Used for the general distinction between row-level and group-aware evaluation, not as evidence about Voynich or as a library used by the standard-library audit code.
Collection: Voynich Research Library · Editorial framework: Wintour House.
EDKSG-VOY-V008 · Edition 1.0.0 · 5 September 2026 · Baseline EDKSG-VOY-B000. Prior public editions are retained. This is a retrospective opportunity and exposure audit, not independent scientific review, complete-transcription validation or decipherment.