VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Voynich | edkSG Vol. 2 — The Input Audit

READER’S GUIDE · Added 5 September 2026 · Original research edition retained

Where this volume fits

This is an input audit, not a translation experiment. It checks the archived matrix and shows how selection rules change the sample; it does not recover the complete token-level corpus.

Previous: Vol. 1 — The Research Baseline. Continue along both checks: Vol. 3 — The Extraction Gate and Vol. 3 — The Source Identity Cross-Check. These are separate published reports sharing a printed volume number.

Voynich | edkSG Research Volumes
Volume: EDKSG-VOY-V002 · Edition: 1.0.0 · Edition date: 5 September 2026, Singapore
Record type: completed matrix audit and prospective test design—not a decipherment or a completed boundary experiment.

Previous: Vol. 1 — The Research Baseline · Voynich Research Library

A sensible-looking instruction can remove an entire part of the manuscript from a study.

Keep only pages with a Currier A or B assignment. Match the writing conditions. Compare like with like. On first reading, this sounds like methodological care.

We applied that selection to the archived matrix—not to the manuscript images—and found that it removes all 20 page units in the M2 class. In this matrix, M2 is the inherited category for radial or circular emblem diagrams. The same selection removes seven of the eleven M4 page units. The resulting dataset is not simply a slightly smaller version of the original. Its composition has changed. [2, 3]

That is the first consequential result of this volume. A control can make an analysis more defensible within a narrower population while making a manuscript-wide conclusion less defensible. We must record both effects.

The audit establishes a usable page-level accounting layer. It does not yet supply the token-level evidence needed to reproduce the inherited boundary claims. This article records what was actually checked, what the checks revealed, and the precise work needed before the next experiment.

In this volume: completed checks · sampling exclusions · next test design · sources and method.

1. The question carried forward from Volume 1

Volume 1 closed the historical baseline without closing its unanswered questions. Its next proposed task was an input and protocol audit for writing units and boundaries, coordinated with the unfinished visual-component investigation. This volume performs the matrix-accounting part of that task and specifies the remaining input requirements. [1]

The question is practical: what can the records presently recovered support us in calculating?

That differs from asking whether Voynich is language, cipher, notation or deliberately structured non-language. It comes before choosing among those possibilities. A claim can have an interesting interpretation and still lack the exact input needed to reproduce its reported statistic.

Three layers must remain distinct here. The manuscript is the historical object. The transcription is a representation of its marks and their locations. Our archived workbook is a further representation: a table of page-level classifications, counts and relationships. Checking the table is worthwhile, but it is not a fresh reading of the parchment.

The audit therefore makes no new identification of a plant, language, workshop or author. It does not complete Test 2A. It does not turn the final-y proposal into a result. Its contribution is to make the next calculation less likely to start from a silently altered sample.

2. What was recovered and examined

The local input was Voynich_Test_1b_Frozen_Segmentation_Matrix.xlsx, the workbook frozen on 2 August 2026. The audit extracted the 227 data rows from its “Frozen Matrix” sheet, retaining the supplied field names and values. An executable audit then checked identities, sequence references, classification coverage, locus-count consistency and the effects of excluding unassigned classifications. [2, 3]

The inspected workbook has this SHA-256 digest:

226b1406bd605f222a033826fb28c05e66db4185d2271e733aed9d7b6c0e415c

The digest identifies the bytes checked in this run. It does not certify that the earlier classifications were correct, that the workbook fully represents the manuscript, or that the original upstream transcription has been recovered. Those are different checks.

The accompanying research record retains the extracted rows, the audit program, its output and the software-test log. The present result is consequently more specific than “the spreadsheet looked consistent”: the checks can be rerun against the retained representation. A repeat using those same rows is a reproducibility check, not an independent palaeographic assessment.

The investigation also inspected the publicly readable ZL and IT transcription headers, the maintained transcription catalogue, and the relevant IVTFF format documentation. Complete raw transcription files were not successfully acquired as verified local byte snapshots during this run. Browser-readable text was therefore used for source identification and format review, not silently substituted for a hash-matched computational corpus. [4–7]

3. The accounting layer holds together

The extracted table contains 227 rows and 227 distinct page identifiers. Its sequence numbers run consecutively from 1 to 227. Each supplied “Previous page” reference agrees with the preceding row, with START at the beginning. The rows use 52 recorded bifolio-group identifiers. [3]

These are checks of the encoded arrangement. They do not prove that the encoded order is the original historical order. A table can consistently record a later arrangement.

CheckObserved result
Data rows / distinct page identifiers227 / 227
Sequence numbersConsecutive, 1–227
Previous-page reference discrepancies0
Recorded bifolio groups52
Geometry summaries disagreeing with their numeric component counts0 of 227 rows
Summed locus-subtype counts disagreeing with their parent counts0 of 227 rows
Completed accounting checks on the archived matrix

The two count-consistency checks are useful because the workbook records the same accounting information in different forms. One field contains numeric counts; another summarises them in a compact geometry signature; another lists subtype totals. Their agreement helps detect extraction and bookkeeping errors. It is not independent confirmation of the observations from which all three fields were derived.

For example, a transcription mistake copied consistently into every summary would survive these checks. To detect it, we would need to return to the transcription or image rather than calculate the same aggregate again.

The result is nonetheless valuable: within the checks performed, the local representation is internally coherent enough to support a transparent coverage audit.

4. What the 5,385 total actually counts

Summing the workbook’s location fields produces the following census. [3]

Recorded location familyCount
P: linear paragraph-text locations4,130
L: short-text or label locations1,029
C: circular-text locations84
R: radial-text locations142
Other0
Total recorded loci5,385
Aggregated text-location counts in the 227-row matrix

The workbook separately totals 740 paragraph starts. They are not another 740 loci to add to the table. They are a different recorded attribute. [3]

A locus need not equal a physical line: IVTFF’s Pt item can share a line with its predecessor. [4, Table 9]

More importantly, this run recovered counts of 5,385 loci. It did not load 5,385 token-bearing records from the original transcription into the analysis. Knowing that a page contains twenty text locations does not reveal the last glyph of its seventh location or the token following a particular separator.

The maintained transcription catalogue reports the same ZL totals of 5,385 loci and 740 paragraphs. This is a useful source-consistency comparison, but the workbook derives from that transcription family. Agreement is expected; it does not provide a second independent witness. [7]

The distinction is the central input finding: page-level counts can audit coverage, but cannot reconstruct token-level transitions.

5. The missing assignments are not evenly distributed

The Currier field contains 114 A assignments, 83 B assignments and 30 unassigned entries, encoded as the literal value NA. Treating NA as a third textual population would misread the workbook. Treating it as though it were an empty spreadsheet cell could also miss it. [3]

We therefore examined what an A/B-only selection retains and removes, using the matrix’s inherited visual categories.

Inherited classOriginal unitsRetainedUnassigned / removed
M11291272
M220020
M319190
M41147
M516160
M632311
Total22719730
Effect of retaining only page units assigned Currier A or B

In the codebook, M2 describes radial or circular emblem diagrams, while M4 describes complex circular, multi-roundel or foldout diagrams. The counts above expose a property of these inherited assignments. They do not independently establish those categories through blind image classification. [2, 3]

Nevertheless, the selection consequence is exact. An A/B-restricted model built from these rows has no M2 observations. It cannot report a measured effect for M2 merely by giving the overall result a manuscript-wide title.

There are two legitimate responses. One is to narrow the claim: the result concerns the retained, assigned population. The other is to design a separate analysis for the unassigned material, with its own comparison conditions. Neither response requires guessing the missing classification.

A third response is not legitimate: silently impute A or B from neighbouring pages, then cite the expanded sample as though the missing assignment had been observed. That would turn a modelling choice into historical evidence.

This finding does not show that an A/B-controlled analysis is wrong. It shows why its population must be stated. The analysis may answer a useful question about 197 page units while leaving a different question about the other thirty unresolved.

6. Two missing clusters and one composite hand

The audit also found that statistical-cluster assignments are absent for fRos and f116v. That is why the cluster coverage is 225 rather than 227. A cluster-dependent calculation needs either to exclude those two page units explicitly or define a separate treatment; it must not make up a cluster. [3]

The header-hand field is unassigned at f115r, while the effective-hand field records 2+3. The latter is not a licence to replace the page with “scribe 2” or “scribe 3” according to convenience. A finer analysis needs the evidence behind the composite assignment. [2, 3]

These details matter because “match by scribe” can conceal different operations. Matching an unambiguous page-level assignment is one operation. Resolving a mixed page to individual locations is another. An input audit should say which operation the available records permit.

The distinction also explains why complete-group denominators differ. Recalculation reproduced 45 complete Currier-classified bifolio groups, all internally uniform on that field, and 50 completely cluster-assigned groups, of which 21 are uniform. These reproduce the inherited matrix summaries under their original coverage conditions. They are not newly independent discoveries about scribal practice. [3]

We retain the unknowns because their locations tell us where a proposed control becomes an assumption.

7. A source name is not an exact computational input

The workbook names ZL3b-n.txt, dated 13 May 2025, as its principal page-and-locus source. Its source sheet also records an expected digest for that upstream file. The publicly inspected ZL header identifies version 3b with that date. [2, 5]

This is a useful identification match. It is not a successful comparison of the complete upstream bytes against the workbook’s expected digest. That comparison remains to be done.

The inspected IT header illustrates why the distinction matters: it identifies version 2a and separately records a modification dated 25 June 2025. A short version label alone does not tell the reader whether two researchers used identical bytes. [6]

The format document has its own numbering: file-format version 2.0.1 and document issue 2.0.2, dated 8 July 2025. These are neither the ZL data version nor the edition number of this article. [4, cover]

For the next computational run, our input record should therefore bind the source, data version, retrieved bytes, format interpretation and analysis settings separately. A filename helps locate a candidate. A content digest helps identify the exact candidate examined. Neither establishes a linguistic reading.

We also need to record ancestry. The maintained catalogue describes IT as a Takahashi-derived interlinear version and RF as an automatically combined reference using ZL and GC. A later comparison must not count a derived reference and its contributing transcription as wholly independent evidence. [7]

The practical decision for this volume is conservative and specific: browser inspection has identified useful source candidates, while byte-level acquisition remains incomplete. No full-corpus token count or boundary effect is reported from those candidates here.

8. Preserve the separator before testing its effect

IVTFF distinguishes a period for a confidently identified gap, a comma for a doubtful small gap, and drawing interruptions represented by <-> or <~>; the latter also marks vertical misalignment. It separately marks paragraph beginnings and endings. These are representation conventions, not decoded punctuation. [4, §6.7]

Our proposed analysis should preserve those distinctions in an untouched source layer, then create separately named analytical views. This is a proposed treatment, not a claim that one interpretation has already won.

Consider the invented expression ab,cd.ef. It is not a manuscript quotation. Joining the doubtful seam yields two analytical tokens, abcd and ef. Splitting it yields three, ab, cd and ef. Excluding the uncertain span leaves a different evaluation sample again.

The disagreement is not just about the word count. It changes which form precedes ef, which endings are counted, and which pairs can enter a transition table. A comparison between analyses must report these changing opportunities, not merely whether both produce an attractive statistic.

For drawing interruptions, our initial boundary protocol should place ordinary adjacent-token comparisons on hold across the interruption. That is a deliberately cautious eligibility rule for this experiment—not proof that a historical reader stopped there. A later, separately labelled analysis can investigate transitions across interruptions.

Likewise, a transcriber’s doubtful-gap marker should not become a measured physical distance. A study of actual gap widths would need image coordinates, scale and measurement procedures. The current workbook does not supply them.

The aim is to keep each result attached to the representation choices that made it calculable.

9. Turn the boundary lead into a precise test

The inherited research record reports context-sensitive choices between forms represented as oR and qoR. It also carries the open test TEXT.SIGMA_Y.001, comparing matched forms with and without final y. These remain provisional structural leads, not translations. [8]

They should not be combined into one vaguely defined “y effect”. The ending of the preceding token and the ending of the current token are different variables. A study must state which one is being compared and which outcome is being predicted.

For the final-y comparison, the first task is to define the eligible pairs. A pair should share an explicitly defined residual transcription string, differing in the target ending under the stated representation. That residual string is an analytical match key; calling it a stem does not establish that it is a linguistic morpheme.

For a prefix-choice comparison, the outcome must likewise be specified: which eligible occurrences count as oR and which as qoR, what qualifies as the same R, and which uncertain readings are excluded or analysed separately. The matching rules must be recoverable before the headline odds ratio is calculated.

The earlier record mentions boundary odds ratios around 5.87 and 6.20. We did not recover the complete original event tables and calculation code needed to reproduce those particular figures in this run. Reprinting the numbers would not close that gap. [8]

The next result should instead expose the denominator: how many eligible occurrences, how many matched pairs, how many exclusions, and which physically related groups contributed them. It should state whether the calculation is intended as a reconstruction of the old result or as a new test under newly specified rules.

This prevents a common form of accidental substitution: a different procedure produces a similar number, and the similarity is announced as a replication.

10. The protocol we can now specify

The following is a prospective protocol direction. It has not been executed on a verified full transcription corpus in this volume, and it is not a claim of external preregistration.

First, fix eligibility. Begin with the declared linear-text population and an explicit treatment of uncertain forms, paragraph openings, ordinary line openings, endings and drawing interruptions. Keep labels and diagram text separate until their own eligibility rules have been justified. Publish the numbers removed by each rule.

Second, fix the contrasts. Keep current-token final y, preceding-token terminal form and oR/qoR prefix choice in separate fields. Define the matched string before examining the outcome. Record where the available matrix supplies only page-level information and where locus-level metadata must be recovered.

Third, compare explanations. A predecessor-based model should be compared with a model using the same available layout and production information but not that predecessor feature. The question is whether the additional feature improves prediction under the stated conditions—not merely whether an association exists in pooled data.

Fourth, preserve the sample boundary. An A/B-only analysis should explicitly report its exclusion of M2 and the other unassigned page units. A separate unassigned-population analysis may be exploratory. It should not quietly inherit the status of the controlled assigned-population test.

Fifth, separate development and evaluation by related material. The existing bifolio identifiers provide a candidate grouping scheme. Before selecting a split, verify their suitability and decide how incomplete or uncertain groups are treated. No definitive held-out set is declared here. Nearby or closely related material must not be divided casually merely to obtain a larger evaluation count.

Finally, report outcomes without semantic inflation. Better prediction would support a bounded structural relationship. No improvement would weaken that particular predictive proposal. Insufficient matched coverage would be an inconclusive result. None of the three outcomes would, by itself, identify the language or assign a meaning to q or y.

11. The visual test has a sampling frame, not a result

The inherited Test 2A starts with the M1 class. The audited matrix contains 129 M1 page units across 33 recorded bifolio groups. Their effective-hand assignments are 95 for hand 1, twenty for hand 2, eight for hand 3 and six for hand 5. [3]

This provides a candidate frame for selecting material. It does not provide 129 independent observations of a proposed visual component. Several page units belong to related physical groups, and the classifications themselves are inherited.

No anonymised crop collection was completed or scored in this volume. The required source-image-to-crop mapping, negative controls and repeat or independent coding remain work to be done. The retrieved Test 2A record remains open. [8]

The next visual preparation should make selection independent of whether an image already looks like the desired answer. It should preserve the source coordinates for later verification while concealing the information that the coding protocol requires to be hidden.

It should also state what “blind” means. Hiding a filename does not erase an analyst’s prior familiarity with a famous folio. Repeat coding by one analyst can test consistency, but is not agreement between independent observers.

The practical advance here is that the population can now be described before choosing crops. That lets us ask whether a proposed batch is disproportionately drawn from one hand or one physical group, instead of discovering the imbalance only after the matches look persuasive.

12. What the software checks do—and do not—validate

The audit program was exercised with ten small software tests. They check specified behaviours including recognition of literal NA, detection of duplicate page identifiers, sequence and previous-reference errors, count inconsistencies, incomplete group denominators, negative counts and preservation of a composite hand value. All ten completed successfully. [3]

These are checks of selected bookkeeping rules, using constructed fixtures. They are not ten tests showing that a manuscript hypothesis is correct. Nor are they a probability that the article contains no errors.

The more important scientific limitation is upstream. We have recomputed a derived matrix. We have not independently reassigned scribal hands, inspected every image, recovered the original text-analysis code, or re-estimated the statistical clusters.

The extraction and the checks were performed within this investigation; no independent reviewer or external replication is claimed. Their useful scope is local and inspectable: the reported counts follow from the retained rows under the recorded rules.

This is enough to support the sampling conclusion. It is not enough to support a decipherment.

13. What changed at this checkpoint

Before this volume, the baseline carried structural summaries and an instruction to recover the inputs for further testing. After this volume, we have an executed local accounting audit, an explicit missing-assignment map, a demonstrated selection consequence and a narrower specification of the data still required.

ItemState at this checkpoint
Archived matrix extraction and accounting checksCompleted within the stated scope
A/B-only sample compositionCalculated; all 20 M2 units excluded
Full upstream transcription byte matchNot established in this run
Original reported boundary-effect reconstructionNot completed; original event tables and code not recovered here
Final-y experimentNot run in this volume
Blind visual-component experimentNot run in this volume
Language, mechanism and plaintextNo new identification
Research-state change recorded by Volume 2

The new audit finding can be referenced locally as EDKSG-VOY-V002-AUDIT-01. The older scientific claims and test identities remain unchanged. An audit record is not a replacement name for an earlier claim.

Volume 1 also remains unchanged. This is the next record in the pathway, not a rewritten account of what the previous volume had already accomplished.

14. The next justified step

The next research job is now specific: recover and retain the exact transcription bytes, compare the ZL digest with the value in the frozen workbook, and validate a location-preserving extraction before measuring any q/y association.

If the digest differs, the next record should identify a new input edition and investigate the difference. It should not force the file to match an old result. If the digest agrees, that closes the byte-identity question—not the question of transcription accuracy or interpretation.

The extraction should then account for accepted and rejected locations, uncertain readings, separators, paragraph context and any finer scribal assignments needed by the test. Only after those counts are checked should the programme freeze the detailed comparison and evaluation procedure.

The visual branch can continue its source-image and sampling preparation separately. It should not borrow confidence from the text audit, and the text branch should not borrow a meaning from an attractive visual resemblance.

We have not learned what the manuscript says in this volume. We have learned which parts of our current representation would disappear under a seemingly careful analysis—and what we must recover before the next statistic can be trusted.

That is a real step along the pathway: not another interpretation placed on the book, but a measurable improvement in how the next interpretation will be tested.

Sources, method and continuation

[1] Predecessor. Voynich | edkSG Vol. 1 — The Research Baseline, edition 1.0.0. The continuation begins with its input-audit task; it does not replace that volume.

[2] Archived input. Voynich_Test_1b_Frozen_Segmentation_Matrix.xlsx, frozen 2 August 2026. “Frozen Matrix”, “Cross-Axis”, “Defeat Ledger” and “Codebook & Sources” are the relevant sheets. The digest above identifies the workbook audited.

[3] This volume’s executed audit. EDKSG-VOY-V002-AUDIT-01, 5 September 2026. Retained materials: extracted matrix-rows.json, executable audit_matrix.py, recomputed output and ten-test log. Counts are computed from the archived matrix, not independently observed from manuscript images.

[4] Format authority. René Zandbergen, IVTFF specification, cover, Table 9 and §6.7. The format definitions are distinct from this article’s proposed analysis rules.

[5] ZL candidate. ZL3b-n, header inspected 5 September 2026. Complete local byte identity was not established.

[6] IT candidate. IT2a-n, header inspected 5 September 2026. No full-corpus calculation is reported from it here.

[7] Source catalogue and ancestry. René Zandbergen, Transliteration of the Text, particularly the source histories and summary table; consulted 5 September 2026.

[8] Inherited test state. voynich-decoder-living-runtime-v0.1.html, version 0.1.0, 3 August 2026: provisional text observations, TEXT.SIGMA_Y.001 and TEST_2A. Neither experiment is reported as completed by this volume.

Background, rather than a competing article: EVA, Transcription and the Segmentation Problem explains the general representation problem. This volume adds the dated audit of the actual inherited matrix.

Catalogue: Voynich Research Library · Editorial framework: Wintour House.

EDKSG-VOY-V002 · Edition 1.0.0 · 5 September 2026 · Predecessor EDKSG-VOY-V001 · Baseline EDKSG-VOY-B000. Later corrections should identify the affected result and edition. This audit does not claim independent replication or decipherment.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading