A method is always smartest on the examples that taught it how to think.
That is exactly why those examples cannot be trusted to grade it.
The Voynich and Padua programme has now accumulated enough known manuscripts to build detailed localisation rules.
That creates the final validation danger:
If we design the rules from the same Paduan, Viennese, German and Bohemian manuscripts on which we measure success, we may only be measuring how well the framework remembers its development set.
The solution is the Out-of-Sample Localisation Test.
Build the rules on one manuscript set.
Freeze them.
Then test them on another set whose answers were not used during design.
Why In-Sample Success Is Not Enough
Suppose Pal. lat. 1311 teaches us that a certain combination of ruling, medical modules and paratext is common in a dated Paduan manuscript.
We add those features to the Padua score.
Then we test Pal. lat. 1311 and recover Padua.
That result is nearly guaranteed to look good.
The interesting question is whether the same frozen feature system can correctly classify a different Paduan manuscript it never saw during rule design—and avoid falsely classifying a new Viennese or German manuscript as Padua.
Prediction Research Calls This Generalisation
In predictive modelling, performance measured on development data is often optimistic. Independent or external validation asks whether the model continues to perform on data not used to construct it.
The same logic applies here even though manuscript localisation is not a clinical prediction model.
We are still building a rule system from examples and asking whether it travels to new examples.
A localisation rule that works only on the manuscripts that inspired it is a description of the training set, not yet a general method.
Development, Calibration and Test Sets Must Have Different Jobs
| Set | Allowed use | Forbidden use |
|---|---|---|
| Development | Discover candidate features, design Coordinate Stack rules, explore weights. | Claim unbiased final performance. |
| Calibration | Adjust thresholds and confidence so scores match known performance. | Repeatedly tune until every calibration object is correct and then call that generalisation. |
| Blind test | Measure a frozen version once against hidden known answers. | Change rules after seeing a result and keep counting the object as blind. |
| External / out-of-sample | Evaluate the final frozen method on fresh manuscripts from the relevant domain. | Use any feature from these objects during method development. |
The Holdout Must Be Chosen Before We Need It
A common mistake is to search for a new test manuscript only after seeing where the model performs badly.
That creates another selection pathway.
The holdout set should be defined before final tuning.
For example, reserve:
- one or more securely catalogued Paduan manuscripts;
- one or more close Viennese controls;
- one or more German or Bohemian controls;
- at least one mixed or uncertain object;
- objects spanning different medical functions so the method cannot rely on one genre shortcut.
Do not open their origin labels during tuning.
Temporal Holdout Is Especially Valuable
One way to make the test harder is to separate manuscripts by date.
Develop on objects around 1400–1425.
Test on objects from 1425–1450.
Or reverse the direction.
This reveals whether the method has learned a region or merely a narrow date-specific style.
For the Voynich, whose parchment window spans decades, temporal robustness matters.
Institutional Holdout Is Another Strong Test
If all development manuscripts come from one modern catalogue project, the method may accidentally learn that catalogue’s descriptive habits.
A stronger test uses independent institutions.
Develop on Bibliotheca Palatina objects.
Test on Wellcome, British Library, Morgan, BnF or other independently catalogued manuscript corpora where suitable.
This is not because one institution is better.
It is because a method should survive changes in catalogue language, photography and collection history.
Geographic Holdout Tests Transportability
Another design is leave-one-region-out validation.
Build the generic feature architecture on all regions except Vienna.
Then test whether the method handles Vienna sensibly when it first encounters it.
Repeat for Padua, Germany or Bohemia.
This is particularly useful for discovering whether our feature vocabulary is secretly region-specific.
Do Not Leak the Holdout Through Article Writing
This project has an unusual problem.
We are publishing the comparator estate while designing the method.
Once a holdout manuscript has been researched in detail, its answer and features are no longer truly unseen to the researchers.
Therefore future formal validation should reserve manuscript objects that have not already been deeply analysed in the Padua branch.
The current published objects can serve development and calibration.
Fresh objects will be needed for credible out-of-sample testing.
Freeze the Entire Pipeline, Not Only the Final Score
Data leakage can occur before scoring.
If we choose features after peeking at the holdout, the holdout has already influenced the model.
So freeze:
- candidate-region definitions;
- feature extraction protocol;
- evidence-family weights;
- dependence discounts;
- expected-evidence rules;
- minimum-path penalties;
- null threshold;
- confidence scale;
- tie-breaking rules.
The holdout should evaluate the whole pipeline.
Out-of-Sample Performance Needs More Than Accuracy
Record several dimensions.
| Metric | Research meaning |
|---|---|
| Correct regional assignment | Did the method recover the independent answer? |
| False Padua rate | Did non-Paduan objects cross the Padua threshold? |
| Null quality | Were ambiguous cases left unresolved rather than forced? |
| Mixed-state quality | Were multi-region objects preserved as complex? |
| Confidence calibration | Were high-confidence calls actually more reliable? |
| Sensitivity | Does the result survive reasonable perturbation of weights and missing features? |
A Failed Holdout Is More Valuable Than a Perfect Development Set
Suppose the frozen model performs beautifully on its development manuscripts and poorly on fresh objects.
That tells us something decisive.
The rules have learned contingent features of the development corpus.
Perhaps the regional categories are too coarse.
Perhaps catalogue metadata leaked into feature design.
Perhaps the localising residue was not actually local.
Perhaps Padua received too much prior weight.
The correct response is not to explain away the holdout.
It is to version the method and begin again.
The Voynich Must Remain Outside the Training Loop
There is a deeper issue.
We ultimately care about Beinecke MS 408.
If every feature and weight is repeatedly adjusted because it makes the Voynich look more or less Paduan, then the unknown target is contaminating the model-development process.
The better discipline is:
develop and validate the localisation method on known manuscripts first; only then freeze a version and apply that version to Voynich.
That would be a genuinely stronger test than tuning the method around the manuscript we hope to localise.
The eduKate Out-of-Sample Protocol
- Define a development corpus of known-origin manuscripts.
- Reserve a holdout corpus before final rule tuning.
- Keep holdout origin labels and detailed analyses unavailable to the scoring team.
- Develop feature extraction and weights only on the development corpus.
- Use calibration objects to set assignment and null thresholds.
- Freeze the complete method version.
- Run blind localisation on the untouched holdout objects.
- Reveal known coordinates once.
- Report correct calls, false positives, nulls, mixed states and confidence calibration.
- If the method is revised, create a new untouched holdout set.
- Only after successful validation apply the frozen method to Beinecke MS 408.
The Validation Ladder Is Now Complete
- Calibration: do confidence and classifications match known outcomes?
- Blinding: can the inference be made without knowing the catalogue answer?
- False positives: how often does the method wrongly identify Padua?
- Out-of-sample: does the frozen method survive manuscripts it did not learn from?
Only after those four steps should a localisation score on Voynich be taken seriously.
Research Sources
- NCBI Bookshelf — Prediction Modeling Methodology — training/test separation, leakage and external validation.
- Prediction models need appropriate internal, internal-external and external validation.
- External validation of machine-learning models — registered models and adaptive sample splitting.
- Statistical analysis of high-dimensional biomedical data — independent test data and over-estimation from development data.
- The Calibration Problem
- The Blind Localisation Test
- The False Positive Problem
- Voynich Research Library
The Final Idea
The method should meet the Voynich only after it has learned to survive without the Voynich.
Build on known books. Freeze the rules. Test on fresh known books. Then—and only then—ask the unknown book where it came from.