Voynich | edkSG Research Volumes · EDKSG-VOY-V011 · Edition 1.0.0 · 5 September 2026, Singapore
Research status: executed retrospective group-transfer and recorded-hand comparison on the existing 685-event collection. No new manuscript observations, prospective validation, independent scientific review or decipherment are claimed.
Leaving out a page does not necessarily leave out its closest recorded neighbours.
Four of the pages evaluated in Volume 10 belong to the same bifolio group in our inherited metadata. When one was removed from fitting, the other three could remain. That does not establish that the previous result was inflated. It identifies a stronger separation that had not yet been tested.
We have now performed that test. Removing all retained observations in the evaluation group leaves six evaluable group folds: four improve and two worsen when the learned predecessor-y coefficient is compared with zero. The total conditional log-likelihood difference is approximately +8.465. The association survives this particular group-separation challenge, but still transfers unevenly. [3]
Then we tried a plausible contextual rule: learn each coefficient only from the same recorded hand, while still excluding the evaluation group. On the observations where both approaches can be compared, this rule performs worse overall. It gives approximately −2.678 against zero, while the common model gives +10.442 on exactly those same observations. One additional hand-specific fit cannot produce a finite estimate and is held rather than forced through. [3]
The new finding is not a meaning for q or y. It is a failed shortcut: the recorded hand alone is not an adequate rule for deciding where to learn the relationship in this collection.
1. Two tests, with their boundaries fixed
The first test changes the separation between fitting and evaluation. Instead of withholding one page, it withholds every retained event assigned to the same recorded bifolio group. The second test changes the training population further: only events with the evaluation hand’s recorded label may contribute to its fitted coefficient.
Both tests were written into the local design record before their new calculations. The underlying results from earlier volumes were already known. This is a retrospective design, not external preregistration and not a recovery of an untouched test set.
The target remains a represented relationship. The current span is either oR or qoR, with R a nonempty remainder. An optional initial q is removed to identify the exact current-form family. The predictor is whether the preceding span ends in y. These are operations on transcription features, not statements about pronunciation, roots, grammar or meaning.
The source representation and event eligibility do not change. We reuse the preserving split-policy indices accumulated through Volume 9. The new calculation does not retokenize text, choose alternative readings, bridge drawing interruptions or admit a previously excluded pair.
That is important for interpretation. A change in transfer performance should come from the declared fitting and separation rules, not from silently changing what counts as an observation. The representation correction in Volume 7 remains part of this experiment’s ancestry. [1, 2]
There is no third step in which we select whichever route looks best on each evaluation page. Such selection would use the outcomes being evaluated to choose the method. We compare two declared procedures and preserve the unfavourable results.
2. One collection, several different denominators
The retained collection contains 685 unique admitted events across fifteen event-contributing page units and nine recorded bifolio groups. The broader exposure ledger includes an additional previously inspected page that contributes no event to this contrast. The present model does not turn that absence into a new observation. [2, 3]
Each event retains its identity, page, location, event-level hand, current-form hash, predecessor class and current outcome. We join its page to the inherited bifolio map. Missing or conflicting identities stop the calculation rather than being guessed.
There are 414 exact page–hand–current-form strata. A stratum is simply a group of observations with those same recorded values. Forty-eight have both current outcomes and contain 222 observations. Of those, 33 also vary in the predecessor class and contain 180 observations. [3]
The distinction between 222 and 180 explains a detail of the transfer score. A stratum in which the predecessor class never changes can contribute a constant to the conditional likelihood, but no information about the predecessor coefficient. Its model-versus-zero difference is zero. The reported conditional evaluation includes it; the stricter informative count does not.
Neither count is the number of independent historical samples. Some neighbouring events share spans. Several pages share recorded production groupings. All the text indices derive from one manuscript and one transcription lineage.
No new source files or images were acquired for this volume. The full origin-transcription reconciliation remains open. This study is reproducible from the retained event indices and metadata, but that does not independently establish the accuracy of their upstream transcription.
3. Recovering Volume 10 before extending it
Before changing the evaluation split, we reconstructed Volume 10’s nine scored page folds using a new implementation of the same one-coefficient conditional likelihood.
The reconstruction preserves the five-improve, four-worsen result. Its summed change is approximately +8.854, compared with +8.854 when the earlier saved result is rounded to three decimals. The largest individual fold difference is below 0.000455, within the numerical tolerance of 0.001 specified in this run’s design. [3]
The small differences arise between the saved numerical optimisation and the new score-root calculation. We do not label that a discovery about the manuscript. The point is that the reconstructed mathematical comparison agrees to the declared tolerance before the grouping rule changes.
The all-data coefficient has an odds-ratio equivalent of approximately 4.914 in the new calculation. That is not the new finding either. It is a reference check against the preceding work.
The new solver uses the counts possible under a binary predictor rather than enumerating every individual allocation during fitting. If a stratum contains n observations, m current qo outcomes and r predecessor-y observations, the possible overlap count ranges from max(0, m−(n−r)) to min(r,m). Combinatorial weights account for how many allocations have each overlap.
We separately tested that calculation against explicit enumeration on small constructed examples and against another conditional-likelihood implementation. Those checks establish agreement between specified computations, not independent evidence about what the manuscript says.
4. What the score does—and does not—predict
The score asks how well a coefficient learned elsewhere accounts for the allocation of qo outcomes within an evaluation stratum, given that stratum’s observed total number of qo outcomes.
That condition is essential. This is not a system receiving an unread page and predicting every token without knowing its outcomes. It uses an observed outcome total to remove the stratum-specific intercept and then compares allocations consistent with that total. Conditional logistic regression is designed to remove group intercepts through conditioning; that does not turn the resulting score into an unconditional prediction. [4]
Our comparison is the conditional log-likelihood with the learned coefficient minus the conditional log-likelihood with the coefficient fixed at zero. Positive means the learned association better supports the observed allocation under this model. Negative means zero does better.
Both alternatives use the same evaluation observations and the same conditioning. The score is neither classification accuracy nor a percentage of decipherment. Its units are natural-log-likelihood units, and its magnitude depends on the amount and arrangement of evaluable information.
This qualification also belongs with Volume 10. The model there was correctly described as conditional, but the word transfer should not be read as a claim of ordinary out-of-sample word prediction. This volume makes that boundary explicit before interpreting another result.
5. Removing the whole recorded group
For each of the nine recorded groups represented in the event collection, we remove all its retained observations from fitting. The page-level conditioning strata remain unchanged. The evaluation is then performed on the omitted group’s outcome-varying strata.
Only six groups contain the variation needed to score the predecessor coefficient. The three others are reported as having no evaluation information for this contrast, not as successes or failures.
| Omitted recorded group | Evaluated page units | Evaluation observations | Change against zero |
|---|---|---|---|
| Q4-B2 | f26r | 9 | −0.535 |
| Q13-B1 | f75r | 86 | +6.317 |
| Q19-B1 | f102r1 | 11 | +1.001 |
| Q19-B2 | f100r, f100v, f101r, f101v | 32 | −1.977 |
| Q20-B2 | f115r | 22 | +0.059 |
| Q20-B4 | f106r | 62 | +3.600 |
The total is +8.465 across six scored group folds. Four improve and two worsen. At page level, the familiar five-improve, four-worsen pattern remains, but the four Q19-B2 pages now share one coefficient trained without any of them.
On Q19-B2, the coefficient learned from the remaining groups has an odds-ratio equivalent of approximately 6.566. It performs worse than zero on the combined conditional evaluation. That makes the group’s resistance to the general positive coefficient visible without letting its other pages influence that fit.
The main change from page exclusion occurs in this four-page group. Most of the other scored groups contribute only one analysed page to the current collection. Whole-group exclusion is therefore a meaningful extension, but not a wholly different experiment for every row.
The positive total remains concentrated. f75r and f106r contribute approximately +9.917 together, more than the net total after the other contributions are included. Removing the same-group training overlap does not make the association uniformly portable.
6. Group separation is not an independence certificate
Every fold records the event identities placed in fitting and evaluation. No retained event from the evaluation bifolio group remains in its training set. That is an inspectable property of the split.
It is not a proof that training and evaluation are historically independent. Group assignments may be imperfect. Copying relationships, handwriting habits, textual families and transcription decisions can extend beyond them. The experiment has removed one declared overlap, not every possible shared cause.
Group-aware evaluation is useful when observations have a known grouping structure, but the right grouping depends on the question. General cross-validation guidance distinguishes such grouped evaluation from random row splitting. It does not decide the manuscript’s physical construction for us. [5]
The nine groups also do not represent a random sample of all manuscript gatherings. They enter through the purposive samples already described in earlier volumes. The complete exposure history is not reset by a stricter split.
Accordingly, the new result remains a retrospective robustness check. We can state that the pattern survives this specific exclusion in aggregate. We cannot state that it has passed an independent prospective test.
7. Does the recorded hand explain where to learn?
Volume 10 suggested that an additional context variable might explain transfer differences. The recorded hand is an obvious candidate, because it is already available in the inherited data and does not need to be invented after seeing a particular failure.
Our test gives that proposal a precise form. For an evaluation hand, remove the entire evaluation group first. Then fit the same single coefficient using only the remaining events assigned that hand. There is no new predictor and no retokenization. The only extra change is the restriction of the training population.
This is deliberately one simple candidate procedure. It is not every possible model involving scribes. A hierarchical model that partially shares information, or a model combining hand with other context, would be a different proposal requiring its own evaluation.
For f115r, the source events carry two hand assignments. They are evaluated separately under the hand-restricted procedure, rather than assigning the whole page whichever hand makes the result easier to describe. Both portions remain outside training when Q20-B2 is evaluated.
A fit must also be mathematically available. No informative training strata produces a hold. A likelihood with no finite maximum produces a hold. We do not silently clip an infinite coefficient, add a penalty or substitute the common model and then call the substituted result a success of hand-specific learning.
8. The proposed shortcut performs worse on the comparable subset
| Evaluation unit | Hand | Common model change | Same-hand model change |
|---|---|---|---|
| f26r | 2 | −0.535 | −1.203 |
| f75r | 2 | +6.317 | −4.457 |
| f102r1 | 1 | +1.001 | +0.132 |
| f115r, hand-2 portion | 2 | −0.865 | −1.778 |
| f115r, hand-3 portion | 3 | +0.923 | +0.943 |
| f106r | 3 | +3.600 | +3.685 |
On these 190 comparable evaluation observations, the same-hand procedure totals −2.678 against zero. The common procedure totals +10.442 on those exact same rows. The paired difference is therefore approximately −13.120 for hand restriction relative to common training. [3]
Four of the six scored hand-specific evaluation units worsen relative to the common procedure; two improve slightly. This is not six independent trials, and the two f115r portions remain parts of one page.
The most conspicuous failure is f75r. Training only on the other retained hand-2 observations yields an odds-ratio equivalent of approximately 0.572. The common group-excluded fit has an equivalent of about 2.853. The first assigns the wrong direction for much of the evaluated alignment and loses approximately 10.774 log-likelihood units relative to the common fit on that page.
The hand-restricted training population behind that f75r coefficient has only three informative strata containing eleven observations. It is not a comprehensive account of hand 2. Its failure could reflect sparse and unrepresentative support, finer contextual differences, or both.
The proper conclusion is therefore procedural: restricting training by recorded hand, in this particular way and on this retained collection, does not improve the transfer comparison overall. It does not show that scribes are irrelevant, that the supplied assignments are wrong, or that no hand-aware model could help.
9. Why one apparently strong fit must remain unreported
For the hand-1 evaluation of Q19-B2, excluding that group leaves only two informative hand-1 training strata, containing six observations. Their arrangement drives the unpenalized likelihood toward an indefinitely large positive coefficient.
There is no finite maximum-likelihood coefficient to carry forward under the declared estimator. The output is HOLD_SEPARATION, not a huge odds ratio presented as strong certainty.
There are legitimate ways to fit sparse separated data, including explicit regularization. None was part of this comparison. Introducing one after encountering the hold would change the estimator and need a separately identified analysis.
The held evaluation contains 32 outcome-varying observations. Those observations remain in the common-model group result but are excluded from the paired hand-versus-common comparison. That is why the common score on the paired subset is +10.442, not the all-group +8.465.
Comparing −2.678 with +8.465 would mix evaluation populations. The paired table avoids that mistake by showing both procedures on the same supported rows.
A hold is not evidence that the manuscript lacks a relationship. It says that this proposed fitting procedure cannot supply the required finite coefficient from this training evidence. Preserving that distinction is part of the result.
10. What the two tests jointly establish
The whole-group result weakens one concern: the positive aggregate transfer score is not eliminated merely by removing the retained same-bifolio neighbours from training. It remains mixed and concentrated, but it does not collapse to zero or become wholly negative under that change.
The hand-restricted result weakens a different proposal: knowing the recorded hand is not sufficient, by itself, to choose a better training subset under the rule tested here. A plausible contextual label has not earned authority to select the model.
Neither result isolates a historical mechanism. The common and hand-restricted models are descriptions of represented outcome allocations. A language process, a copying process, state-dependent generation, remaining representation choices or several interacting factors could be compatible with parts of the evidence.
The tests also do not locate the missing explanatory variable. Choosing another metadata field because the hand-only rule failed would be a new hypothesis, not an automatic next stage of confirmation. Testing many alternatives on the same exposed pages would require an honest record of that model search.
What we have gained is a constraint on the next model: it must justify why its context division supplies transferable evidence, rather than assuming that a familiar classification necessarily does so.
11. The checks behind this report
The executable analysis uses the Python standard library. It verifies the recorded SHA-256 identities of its three parent inputs, rejects duplicate event identities and inconsistent page/location or hand/frame assignments, and preserves the primary split policy.
Twenty-eight unit and regression tests passed. They include exact small-case allocation checks, conditional-likelihood arithmetic, gradient checks, finite-estimate handling, positive and negative separation, duplicate rejection, metadata consistency and fold disjointness. These are software tests, not twenty-eight scientific experiments.
For a separate numerical comparison, 35 supported training configurations were evaluated with statsmodels’ conditional-likelihood implementation. Their log-likelihoods agreed with the new count-based calculation to a maximum absolute difference of approximately 1.5 × 10−14. The comparison implementation shares the same admitted inputs; it is not an independent transcription or external review. [3, 4]
The archive includes the exact parent analytical indices, normalized events, all strata, fitting/evaluation membership for every group fold, page contributions, hand-specific holds, test output and the Volume 10 numerical reconciliation. It does not require network access to reproduce the primary calculations.
The analysis is reproducible at that layer. It has not re-examined the original images, verified every earlier transcription decision, or closed the full-source acquisition gap. More precise computation does not remove those upstream limits.
12. What the next research should earn
The next useful model should not simply attach the word context to an unrestricted search. It needs a small declared set of explanations, enough training variation to estimate them, and evaluation material not used to choose among them.
The immediate design lesson is to distinguish context recognition from model selection. A recorded hand can be a valid descriptive attribute while being a poor rule for restricting the data used to learn this coefficient. A future model must show what that attribute adds, on an equal evaluation population, without hiding the cases where support is missing.
Broader verified source coverage remains important because this collection supplies very uneven within-context comparisons. Exact acquisition and reconciliation of the complete intended transcription would support a more complete opportunity map. It would not, by itself, make that corpus independent validation material.
All pages in this volume are now explicitly development-exposed. A later prospective test needs a separately justified exposure history, fixed representation and frozen comparison. Current-final-y and visual-component investigations remain distinct open tasks, not results supplied by this transfer analysis.
The group test preserves a limited structural clue. The hand test rejects a convenient way of routing it. The next explanation has to account for both.
Sources and edition record
[1] Published predecessor. Vol. 10 — The Transfer Gate, WordPress post 154202, edition 1.0.0. Its saved page-fold results supply the numerical reference, not new independent observations.
[2] Analytical inputs and grouping. The retained packages for Vol. 9 — The New-Context Check and Vol. 8 — The Comparison-Opportunity Map. The 379 earlier compatible events, 306-event extension and inherited 227-unit frame are identified by hash. The representation qualification in Vol. 7 remains in force.
[3] Executed local record. EDKSG-VOY-V011-GROUP-TRANSFER-01. The accompanying package contains the design, source bindings, standard-library analysis, event and fold ledgers, all comparison outcomes, tests and numerical checks. No fresh source collection, independent scientific review or deployment is asserted.
[4] Conditional-model reference. statsmodels: ConditionalLogit, consulted 5 September 2026. Used for the model’s conditioning interpretation and as a separate numerical likelihood implementation, not as evidence about Voynich.
[5] Evaluation reference. scikit-learn: Cross-validation, evaluating estimator performance, particularly grouped evaluation and separation of model selection from final testing; consulted 5 September 2026. The manuscript-specific grouping remains our inherited analytical choice.
Catalogue continuity: Vol. 3A is The Extraction Gate; Vol. 3B is The Source Identity Cross-Check. Their original URLs and historical identifiers are preserved. This volume follows Volume 10; no earlier work is renumbered or rewritten.
Collection: Voynich Research Library · Editorial framework: Wintour House.
EDKSG-VOY-V011 · Edition 1.0.0 · 5 September 2026 · Baseline EDKSG-VOY-B000. Retrospective conditional-transfer comparison. Earlier editions remain preserved. No universal rule, independent review, prospective validation or semantic decipherment is claimed.