Voynich | edkSG Research Volumes · EDKSG-VOY-V012 · Edition 1.0.0 · 5 September 2026, Singapore
Research status: executed retrospective partial-pooling and nested-selection comparison on the existing 685-event collection. This is a new estimator applied to already-exposed evidence, not new manuscript observations, independent review, prospective validation or decipherment.
A model can become easier to estimate without becoming better at explaining a different part of the manuscript.
In Volume 11, the same-hand procedure encountered an infinite-estimate problem. In this volume, an explicitly regularized procedure produces finite coefficients for every fitted configuration. That repairs a mathematical difficulty. It does not, by itself, repair the proposed explanation.
Our primary partial-sharing model scores +2.052 against a zero-effect reference across the six evaluable recorded groups. A simpler common model, given the same baseline regularization and scored on the same 222 observations, scores +8.553. The more flexible model therefore loses approximately 6.501 conditional log-likelihood units relative to that comparator. [2]
We also let an inner, training-only comparison choose how much sharing to use. It chooses the common model in five folds and partial sharing in one. Its total score is +3.852, still below the fixed common comparator. The single different choice occurs when f75r is the outer evaluation page; there, the selected partial-sharing rule performs substantially worse than the common rule. [2]
The lesson is not that partial pooling is generally wrong. It is that neither a finite estimate nor a more elaborate selection procedure supplies the missing evidence that recorded-hand differences improve this particular transfer task.
1. A middle position between one rule and isolated rules
The previous experiment tested two extremes. A common coefficient learned the represented predecessor-y association from all available training hands. The hand-only procedure learned separate coefficients from each recorded hand, refusing to borrow observations from the others.
The second extreme was vulnerable to sparse training. When f75r was excluded with its recorded bifolio group, only three informative same-hand training strata, containing eleven observations, remained. Another hand-specific training set produced separation: its unpenalized likelihood had no finite maximum. Those were findings of Volume 11, not newly discovered facts here. [1]
A natural next proposal is partial pooling. Let the hands differ, but discourage large differences unless the admitted data support them. This sits between complete pooling and entirely separate estimation. The general distinction is described in the Stan project’s methodological guidance; it is not a new principle invented for Voynich. [4]
We test one particular implementation: a common coefficient plus a penalized hand-specific deviation. The deviations can be positive or negative. They are not forced to strengthen the original clue.
This is a new, declared estimator. We do not retrospectively change Volume 11’s hold into a successful unpenalized fit. Its estimator had no finite answer in that case. Today’s estimator has a finite answer because today’s mathematical objective includes an additional assumption.
The question is whether that assumption earns its place through a better comparison outside the group used for fitting—not merely whether the software returns a number.
2. The evidence is unchanged
The input remains 685 unique admitted transitions across fifteen contributing page units and nine recorded bifolio groups. The normalized event file is copied from Volume 11 and bound to its SHA-256 identity. The source population, separation of text locations, doubtful-gap treatment and plain-span eligibility rules are not changed. [1, 2]
Each event records whether the preceding represented span ends in y and whether the current represented span is qoR rather than oR. Removing an optional initial q identifies the current-form family for matching. R is simply a nonempty remainder in a transcription string, not a decoded root.
We retain exact page, event-level hand and current-form family as the conditioning strata. There are 414 such strata. Forty-eight contain both current outcomes and supply 222 conditional evaluation observations. Thirty-three also vary in the predecessor class and supply 180 directly informative observations. [2]
Those denominators do different jobs. An outcome-varying stratum with no predecessor variation contributes a constant to the conditional likelihood and zero to the model-versus-zero difference. Its observations are included in the 222, but not in the 180.
No new source transcription, glyph reading or manuscript image enters this study. The hand and bifolio assignments remain inherited metadata. A fresh attempt to retrieve the complete origin-host transcription failed at name resolution; that attempt is recorded separately and contributes no data to the model.
The research advance is therefore a tested model comparison on retained evidence. It must not be counted as a new historical witness or as growth in the manuscript sample.
3. The additional assumption is explicit
Write the coefficient used for hand h as β + δh. β is the common component; δh is that hand’s deviation. The coefficient concerns the association with the preceding y feature, not the meaning of a glyph.
The fitting objective is the total conditional log-likelihood minus two penalties:
sum of conditional log-likelihoods
− (0.25 / 2) × β²
− (λ / 2) × sum of δ_h²
The first penalty is applied to the common component in every newly compared model. The second controls how strongly hand deviations are pulled toward zero. A larger λ favours stronger sharing. A smaller λ allows more hand-specific variation.
The primary comparison fixes λ = 1. The sensitivity grid, declared before the new outcome calculations, also contains 0.25, 4 and 16. The common comparator sets every deviation to zero while retaining the same 0.25 penalty on β. We report Volume 11’s unpenalized common result separately.
These numerical settings are analytical choices, not parameters calibrated from an independent Voynich experiment. The penalty applies to the sum of log-likelihoods, not their per-observation average. That scale matters when someone reproduces the calculation.
Positive quadratic penalties make the declared objective strictly concave in its free parameters and prevent coefficient estimates from escaping to infinity. This follows from the objective: the conditional log-likelihood is concave, while the penalties supply strictly negative curvature. It does not imply that the observations contain strong information about each parameter.
This is a penalized point-estimation procedure. It is not a full Bayesian posterior analysis, and we do not report posterior probabilities, credible intervals or a data-estimated population distribution of scribes. A hand absent from informative training receives a zero deviation under the declared rule, with that absence recorded.
4. The evaluation remains conditional
For every outer fold, all events assigned to the evaluation bifolio group are excluded from fitting. The common component and the hand deviations are learned from the remaining groups and then held fixed.
The evaluation asks how well those coefficients account for the observed allocation of qo outcomes within each exact-form stratum, given the stratum’s observed total number of qo outcomes. Conditional logistic modelling removes the stratum intercept through conditioning; it does not supply an unconditional prediction of an unread token. [3]
We subtract the conditional log-likelihood with every coefficient fixed at zero from the conditional log-likelihood using the fitted coefficients. Positive differences favour the fitted association; negative differences favour zero. Penalties are used only during fitting and are not subtracted from the evaluation score.
All newly compared models use the same 222 outcome-varying evaluation observations across the six informative outer groups. Three other recorded groups have no evaluation information for this contrast and remain unscored—not classified as failures.
The result is not accuracy, a percentage of translated words or a probability that the historical hypothesis is correct. Scores depend on the number and arrangement of evaluable observations, so six group folds must not be treated as six equally informative votes.
Withholding whole recorded groups removes one specified overlap. It does not prove historical independence, correct every inherited grouping or make previously studied pages unseen again.
5. The primary partial-sharing model loses on five of six groups
| Omitted recorded group | Evaluation page units | Common ridge score | Partial-sharing score, λ = 1 |
|---|---|---|---|
| Q4-B2 | f26r | −0.508 | −0.909 |
| Q13-B1 | f75r | +6.137 | +1.436 |
| Q19-B1 | f102r1 | +0.984 | +0.521 |
| Q19-B2 | f100r, f100v, f101r, f101v | −1.820 | −2.226 |
| Q20-B2 | f115r | +0.113 | −0.454 |
| Q20-B4 | f106r | +3.648 | +3.685 |
| Total | Same 222 observations | +8.553 | +2.052 |
The partial-sharing model improves over the common comparator only on Q20-B4, by about 0.036 units. It performs worse on the other five groups. Its total remains above zero, but that is not enough to justify its added flexibility when a simpler specified comparator performs substantially better.
The two benchmarks must stay separate. The unpenalized common procedure from Volume 11 scores +8.465 when reproduced here. Today’s common ridge procedure scores +8.553. The small difference reflects a changed estimator, not additional evidence. The primary comparison uses the latter so both new models share the same baseline penalty.
The largest gap occurs at f75r. Partial sharing learns a hand-2 coefficient with an odds-ratio equivalent of about 1.224, compared with about 2.746 for the common ridge coefficient learned without that group. Its conditional evaluation is consequently much weaker. [2]
This does improve on the direction of Volume 11’s isolated hand-only coefficient, which had an odds-ratio equivalent below one. But improving a failed specialist estimate is not the same as beating the common estimate. The relevant comparison cannot stop at the most flattering baseline.
6. A finite answer is not a successful explanation
Q19-B2 illustrates the distinction particularly clearly. When that group is omitted, the remaining hand-1 evidence contains only two informative strata with six observations. Under Volume 11’s unpenalized same-hand procedure, their arrangement produced separation and no finite fitted coefficient. [1]
The partial-sharing procedure now produces a finite hand-1 coefficient, equivalent to an odds ratio of approximately 7.334. The newly available number is a consequence of the regularized model borrowing information and limiting the parameter. It is not evidence that six observations suddenly became a reliable, comprehensive description of hand 1.
On the excluded group’s 32 outcome-varying observations, that partial-sharing fit scores −2.226 against zero. The common ridge fit scores −1.820 on the same observations. Both are unfavourable, and the more flexible fit is worse. [2]
The computational problem and the research problem have therefore diverged. We have a stable finite estimate, while the transfer evidence still does not support using it as an improved explanation of the omitted group.
This is why the earlier hold remains historically valid. It described the limits of an estimator without regularization. The new estimator changes those limits by assumption. A catalogue should record that change rather than retroactively turning the old hold into proof that the original procedure worked.
7. Stronger sharing helps here, but the common comparator still leads
| Declared procedure | Total change against zero | Groups improving / worsening |
|---|---|---|
| Unpenalized common, Volume 11 reference | +8.465 | 4 / 2 |
| Common ridge, no hand deviations | +8.553 | 4 / 2 |
| Partial sharing, λ = 16 | +7.810 | 4 / 2 |
| Partial sharing, λ = 4 | +6.024 | 3 / 3 |
| Partial sharing, λ = 1 | +2.052 | 3 / 3 |
| Partial sharing, λ = 0.25 | −2.693 | 2 / 4 |
Within this declared grid, increasing the constraint on hand deviations brings the total closer to the common model’s performance. None of the four partial-sharing settings overtakes that comparator.
This is a result about the selected grid and sample, not a theorem that every intermediate penalty would fail. We have not searched an unlimited continuum, varied all baseline penalties or tested every possible hand-aware model. Doing so after inspecting the table would create a larger model-search history, not independent confirmation.
The weaker-sharing setting, λ = 0.25, is worse even than the zero-effect reference in aggregate. All its coefficients are nevertheless finite. The ability to calculate them is therefore a poor substitute for evaluating them.
The table supports a restrained decision: these particular hand deviations have not earned preference over the common comparator. It does not establish that the common model is historically correct or that a recorded hand can never help a different, better-supported question.
8. Can training-only selection choose the right amount of sharing?
The fixed grid leaves a further question: perhaps each outer training set should choose the amount of pooling for itself.
We tested that possibility without using the outer evaluation outcomes to choose the candidate. For each scored outer group, the remaining groups undergo an inner leave-one-group-out comparison. Each candidate is fitted without both the outer group and the current inner evaluation group. Its inner conditional scores are then summed.
The candidate set contains the common ridge model and the four declared partial-sharing penalties. Ties within 10−10 favour the common model, then the stronger-sharing option. The candidate selected by the inner comparison is refitted on the complete outer training set and scored once on the outer evaluation group.
This follows the general nested-evaluation distinction between parameter selection and final scoring. The official scikit-learn example explains why using the same results for both can make scores overly optimistic. Our implementation uses recorded manuscript groups rather than the example’s dataset, and does not invoke scikit-learn to perform the experiment. [5]
The procedure still cannot erase our prior knowledge of these pages. Nested separation protects a boundary inside the present computation. It is not a certificate that the model family, candidate grid or research question was chosen before any relevant evidence had been seen.
The outer test remains retrospective, with the complete sequence of prior volumes part of its development history.
9. The selector makes one different choice—and loses there
The inner comparison selects the common ridge model in five of the six scored outer folds. It selects partial sharing with λ = 1 only when Q13-B1, represented here by f75r, is excluded as the outer evaluation group. [2]
Without f75r, the summed inner scores favour that partial-sharing candidate: approximately +3.281, compared with +2.612 for the common ridge candidate. That is the evidence available to the selector under its declared rule.
But on f75r itself, the selected procedure scores +1.436, while the fixed common ridge procedure scores +6.137. The selector therefore loses approximately 4.701 on the only outer fold where it changes the choice. Its total is +3.852 rather than +8.553. [2]
This is not a case of the outer outcomes secretly choosing the model. The inner selection history is saved, and the outer group is excluded throughout it. It is an example of a legitimate training-only selection procedure choosing a model that transfers less successfully to a different retained context.
Nested evaluation is useful because it reveals that failure. It does not promise that a selector will always outperform a fixed comparator, especially when the available contexts differ and the number of informative groups is small.
We do not replace the chosen model after seeing f75r. We also do not report the best outer result from each candidate as though an implementable rule had selected it. Either action would answer a more favourable question than the one we actually tested.
10. What the numerical checks establish
The input identities and inherited likelihood implementation were verified before execution. Reconstructing Volume 11’s common group folds recovers its total +8.4647668164 within the declared tolerance of 10−8. The new calculation preserves every outer group, including the three with no informative evaluation.
Thirty-four new unit and regression tests passed. They exercise source-hash rejection, coefficient constraints, exact small-case likelihoods, gradient and Hessian calculations, finite regularized separation, hand-label invariance, missing-information handling, training-group exclusions, equal evaluation populations and deterministic candidate selection. The unchanged parent suite’s 28 tests also passed. [2]
A second numerical path uses statsmodels’ conditional-likelihood implementation. Across 210 fitted configurations, its penalized objectives agree with the new implementation to a maximum absolute difference of approximately 5.7 × 10−14. A separate optimizer was also run on all 45 outer candidate fits; its largest parameter difference is approximately 2.7 × 10−8. [2]
Twelve of those optimizer runs returned a precision-loss warning rather than a success flag. We retain the warnings. Their objective and parameter differences nevertheless met the numerical acceptance tolerances stated in the checker. We do not describe the warning flags as successful convergence or omit them because the numbers agree.
These checks support the implementation and reproducibility of the stated calculation. Both paths use the same admitted data and were prepared in this investigation. They are not independent scholarly review, new manuscript trials or proof that the representation captures the original writing correctly.
The fitting procedure needs NumPy; the separate numerical check also needs SciPy and statsmodels. The package records the installed versions and includes the inputs, fixed design, code, fitted parameters, inner selections and group-membership ledger.
11. What this changes in the working explanation
We can now distinguish three failures that should not be merged.
Insufficient support for an unpenalized estimate: Volume 11 identified a hand-specific training case with separation. Today’s penalty supplies a finite answer, but that answer includes a modelling assumption.
Insufficient transfer benefit from hand deviations: the primary regularized partial-sharing model loses to the common comparator on five of six scored groups. That is an empirical result within the retained analytical collection, not a conclusion from the penalty formula alone.
Insufficient reliability of the tested selection rule: the nested selector chooses the less successful candidate on its one differing outer fold. The evaluation boundary is preserved, and the choice still fails to improve the total.
None proves that the manuscript has one writing mechanism, that scribes do not matter, or that the predecessor-y association is meaningless. The pooled and common conditional results remain positive in aggregate. What has not been established is the extra value of this hand-specific elaboration.
The general linguistic, encoding, copying and structured-generation possibilities remain unresolved. A statistical coefficient, even one estimated stably and evaluated carefully, does not assign a meaning to q, y or the spans built around them.
The current-final-y and blind visual-component experiments remain separate. No success or failure in this model comparison completes either task.
12. What the next step should add
The immediate research lesson is not to search indefinitely for a penalty that makes these already-exposed pages look convincing. That would increase model flexibility faster than it increases evidence.
A more useful next acquisition should broaden the contexts capable of supporting direct same-form comparisons, or supply measurements that distinguish why the common relationship transfers well in some places and poorly in others. Those measurements need their own source authority; they cannot be invented from whichever model residual looks most striking.
Complete source reconciliation remains valuable because our current sample is uneven and limited. It would make a wider opportunity census possible under one preserved procedure. It would not automatically turn that wider corpus into independent validation material.
For a future predictive study, the source edition, representation, model family, selection rule and genuinely reserved evaluation material must be specified together. The pages examined here remain development-exposed, regardless of how many times they are left out of a particular fit.
Volume 12 repairs a finite-estimation problem but declines to promote the repaired model. The common structural clue remains; the proposed hand-specific improvement has not earned its place.
That is a useful stopping point for this model branch and a clearer entrance to the next evidence question.
Sources and reproducibility record
[1] Exact predecessor. Vol. 11 — The Group-and-Hand Transfer Check, WordPress post 154243, edition 1.0.0. Its normalized events, group results and conditional-likelihood implementation are retained unchanged as identified parent inputs.
[2] Executed local record. EDKSG-VOY-V012-PARTIAL-POOLING-01. The accompanying research package contains the pre-execution retrospective design, hashes, 685-event input, five-candidate grid, all group folds, inner-selection results, fit parameters, source-access outcome, test logs and second-implementation checks. No new transcription, source observation, independent review or operational model activation is asserted.
[3] Conditional-model reference. statsmodels: ConditionalLogit, consulted 5 September 2026. Used for the conditioning interpretation and second likelihood implementation, not as evidence about Voynich.
[4] Partial-pooling reference. Stan: Hierarchical Partial Pooling for Repeated Binary Trials. Used for the general distinction between common, separate and partially shared estimation. The present penalized conditional model, its penalty settings and its evidence are separately declared; no Stan model was fitted in this run.
[5] Selection and evaluation reference. scikit-learn: Nested versus non-nested cross-validation, consulted 5 September 2026. Used for the general separation of tuning and evaluation. The manuscript-group implementation is part of this research package.
Catalogue continuity: Vol. 3A — The Extraction Gate and Vol. 3B — The Source Identity Cross-Check retain their original URLs and historical identifiers. Volume 7’s representation qualification remains in force; no previous research edition is rewritten here.
Collection: Voynich Research Library · Editorial framework: Wintour House.
EDKSG-VOY-V012 · Edition 1.0.0 · 5 September 2026 · Baseline EDKSG-VOY-B000. Retrospective penalized conditional-transfer comparison. No independent scientific review, prospective validation, calibrated probability or semantic decipherment is claimed.