VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich and Padua | The Out-of-Sample Localisation Test: Build the Rules on One Manuscript Set, Then Test Them on Another

A method is always smartest on the examples that taught it how to think.

That is exactly why those examples cannot be trusted to grade it.

The Voynich and Padua programme has now accumulated enough known manuscripts to build detailed localisation rules.

That creates the final validation danger:

If we design the rules from the same Paduan, Viennese, German and Bohemian manuscripts on which we measure success, we may only be measuring how well the framework remembers its development set.

The solution is the Out-of-Sample Localisation Test.

Build the rules on one manuscript set.

Freeze them.

Then test them on another set whose answers were not used during design.


Why In-Sample Success Is Not Enough

Suppose Pal. lat. 1311 teaches us that a certain combination of ruling, medical modules and paratext is common in a dated Paduan manuscript.

We add those features to the Padua score.

Then we test Pal. lat. 1311 and recover Padua.

That result is nearly guaranteed to look good.

The interesting question is whether the same frozen feature system can correctly classify a different Paduan manuscript it never saw during rule design—and avoid falsely classifying a new Viennese or German manuscript as Padua.

Prediction Research Calls This Generalisation

In predictive modelling, performance measured on development data is often optimistic. Independent or external validation asks whether the model continues to perform on data not used to construct it.

The same logic applies here even though manuscript localisation is not a clinical prediction model.

We are still building a rule system from examples and asking whether it travels to new examples.

A localisation rule that works only on the manuscripts that inspired it is a description of the training set, not yet a general method.

Development, Calibration and Test Sets Must Have Different Jobs

SetAllowed useForbidden use
DevelopmentDiscover candidate features, design Coordinate Stack rules, explore weights.Claim unbiased final performance.
CalibrationAdjust thresholds and confidence so scores match known performance.Repeatedly tune until every calibration object is correct and then call that generalisation.
Blind testMeasure a frozen version once against hidden known answers.Change rules after seeing a result and keep counting the object as blind.
External / out-of-sampleEvaluate the final frozen method on fresh manuscripts from the relevant domain.Use any feature from these objects during method development.

The Holdout Must Be Chosen Before We Need It

A common mistake is to search for a new test manuscript only after seeing where the model performs badly.

That creates another selection pathway.

The holdout set should be defined before final tuning.

For example, reserve:

  • one or more securely catalogued Paduan manuscripts;
  • one or more close Viennese controls;
  • one or more German or Bohemian controls;
  • at least one mixed or uncertain object;
  • objects spanning different medical functions so the method cannot rely on one genre shortcut.

Do not open their origin labels during tuning.

Temporal Holdout Is Especially Valuable

One way to make the test harder is to separate manuscripts by date.

Develop on objects around 1400–1425.

Test on objects from 1425–1450.

Or reverse the direction.

This reveals whether the method has learned a region or merely a narrow date-specific style.

For the Voynich, whose parchment window spans decades, temporal robustness matters.

Institutional Holdout Is Another Strong Test

If all development manuscripts come from one modern catalogue project, the method may accidentally learn that catalogue’s descriptive habits.

A stronger test uses independent institutions.

Develop on Bibliotheca Palatina objects.

Test on Wellcome, British Library, Morgan, BnF or other independently catalogued manuscript corpora where suitable.

This is not because one institution is better.

It is because a method should survive changes in catalogue language, photography and collection history.

Geographic Holdout Tests Transportability

Another design is leave-one-region-out validation.

Build the generic feature architecture on all regions except Vienna.

Then test whether the method handles Vienna sensibly when it first encounters it.

Repeat for Padua, Germany or Bohemia.

This is particularly useful for discovering whether our feature vocabulary is secretly region-specific.

Do Not Leak the Holdout Through Article Writing

This project has an unusual problem.

We are publishing the comparator estate while designing the method.

Once a holdout manuscript has been researched in detail, its answer and features are no longer truly unseen to the researchers.

Therefore future formal validation should reserve manuscript objects that have not already been deeply analysed in the Padua branch.

The current published objects can serve development and calibration.

Fresh objects will be needed for credible out-of-sample testing.

Freeze the Entire Pipeline, Not Only the Final Score

Data leakage can occur before scoring.

If we choose features after peeking at the holdout, the holdout has already influenced the model.

So freeze:

  • candidate-region definitions;
  • feature extraction protocol;
  • evidence-family weights;
  • dependence discounts;
  • expected-evidence rules;
  • minimum-path penalties;
  • null threshold;
  • confidence scale;
  • tie-breaking rules.

The holdout should evaluate the whole pipeline.

Out-of-Sample Performance Needs More Than Accuracy

Record several dimensions.

MetricResearch meaning
Correct regional assignmentDid the method recover the independent answer?
False Padua rateDid non-Paduan objects cross the Padua threshold?
Null qualityWere ambiguous cases left unresolved rather than forced?
Mixed-state qualityWere multi-region objects preserved as complex?
Confidence calibrationWere high-confidence calls actually more reliable?
SensitivityDoes the result survive reasonable perturbation of weights and missing features?

A Failed Holdout Is More Valuable Than a Perfect Development Set

Suppose the frozen model performs beautifully on its development manuscripts and poorly on fresh objects.

That tells us something decisive.

The rules have learned contingent features of the development corpus.

Perhaps the regional categories are too coarse.

Perhaps catalogue metadata leaked into feature design.

Perhaps the localising residue was not actually local.

Perhaps Padua received too much prior weight.

The correct response is not to explain away the holdout.

It is to version the method and begin again.

The Voynich Must Remain Outside the Training Loop

There is a deeper issue.

We ultimately care about Beinecke MS 408.

If every feature and weight is repeatedly adjusted because it makes the Voynich look more or less Paduan, then the unknown target is contaminating the model-development process.

The better discipline is:

develop and validate the localisation method on known manuscripts first; only then freeze a version and apply that version to Voynich.

That would be a genuinely stronger test than tuning the method around the manuscript we hope to localise.

The eduKate Out-of-Sample Protocol

  1. Define a development corpus of known-origin manuscripts.
  2. Reserve a holdout corpus before final rule tuning.
  3. Keep holdout origin labels and detailed analyses unavailable to the scoring team.
  4. Develop feature extraction and weights only on the development corpus.
  5. Use calibration objects to set assignment and null thresholds.
  6. Freeze the complete method version.
  7. Run blind localisation on the untouched holdout objects.
  8. Reveal known coordinates once.
  9. Report correct calls, false positives, nulls, mixed states and confidence calibration.
  10. If the method is revised, create a new untouched holdout set.
  11. Only after successful validation apply the frozen method to Beinecke MS 408.

The Validation Ladder Is Now Complete

  • Calibration: do confidence and classifications match known outcomes?
  • Blinding: can the inference be made without knowing the catalogue answer?
  • False positives: how often does the method wrongly identify Padua?
  • Out-of-sample: does the frozen method survive manuscripts it did not learn from?

Only after those four steps should a localisation score on Voynich be taken seriously.

Research Sources

The Final Idea

The method should meet the Voynich only after it has learned to survive without the Voynich.

Build on known books. Freeze the rules. Test on fresh known books. Then—and only then—ask the unknown book where it came from.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading