Voynich | edkSG Research Volumes
Volume: EDKSG-VOY-V003 · Edition: 1.0.0 · Edition date: 5 September 2026, Singapore
Record type: source-identity cross-check and provenance repair—not a decipherment, semantic result, completed q/y experiment or completed visual test.
Previous: Vol. 2 — The Input Audit · Vol. 1 — The Research Baseline · Voynich Research Library
Volume 2 ended with a deliberately small next job: identify the exact transcription source behind the inherited matrix before measuring another boundary effect. The instruction was simple. Recover the source identity. Compare the digest. Do not force a mismatch to agree with an old result. Do not confuse a matching hash with a decipherment.
This volume records what we can now close—and what remains open.
The frozen Voynich_Test_1b_Frozen_Segmentation_Matrix.xlsx records the Zandbergen–Landini source as ZL3b-n, 2025-05-13 and binds it to the following expected SHA-256:
bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc
A separately maintained public Voynich-data preservation project records the exact same full digest for ZL3b-n.txt, identifies the file as Zandbergen–Landini version 3b from May 2025, records a size of 411,671 bytes, and reports the source as the voynich.nu ZL3b file. Its checksum ledger repeats that exact digest beside the filename. [2–4]
A second public preservation route, the cipher-benchmark repository, also preserves a ZL3b snapshot, identifies the same upstream URL, records the version header as Version 3b of 13/05/2025, and documents its transformation choices separately from the source. [5–6]
The narrow result is therefore real: the source identity frozen into our Test 1b matrix agrees exactly, at the SHA-256 level, with a separately preserved public ZL3b verification record.
That result does not mean we have independently re-downloaded and hashed the bytes currently served by the primary origin today. It does not certify every transcription decision. It does not prove that every glyph grouping corresponds to a historical character. It does not recover the original event tables behind the inherited q/y statistics. And it does not move the Voynich Manuscript one step closer to plaintext by itself.
It closes a smaller and necessary uncertainty: the frozen analytical matrix was bound to a source identity that can now be independently cross-checked against a preserved public checksum record.
Current Wintour gate: source identity cross-check PASS within scope; current primary-host byte identity HOLD; transcription accuracy HOLD at the level not independently examined; q/y test HOLD; Test 2A HOLD; semantics HOLD; decipherment HOLD.
1. Why a Source Hash Deserves Its Own Volume
It would be easy to dismiss this checkpoint as plumbing. We already had a filename. We already had a version. We already had a spreadsheet built from the transcription. Why stop to write about a hash?
Because a filename is not a dataset.
Two files can share a filename and differ by one corrected line, one changed metadata field, one altered uncertain reading or one regenerated export. A human reader may consider the difference trivial. A boundary statistic can be sensitive to exactly the place where the difference occurred.
A version label is stronger, but it is still not exact byte identity. “Version 3b” describes an edition family. It does not, by itself, establish that two analysts possess identical copies of that edition.
A cryptographic digest performs a narrower job. For a specified byte sequence and algorithm, it supplies a compact content identity. If the expected SHA-256 is the same and the bytes genuinely produce that digest, the files can be treated as the same byte-level object for that comparison.
That is not the same as saying the file is correct.
A perfectly preserved transcription can preserve a transcription mistake perfectly. A hash protects identity, not truth. It tells us which representation we are talking about. It does not tell us whether the representation captured every physical mark correctly or whether its segmentation corresponds to the manuscript’s historical writing system.
This distinction is exactly why the new-generation Voynich volumes are chronological research records rather than another generation of thematic speculation. Before we ask a representation to support a finer inference, we make sure we know which representation it is.
2. The Frozen Matrix Already Contained a Source Lock
The archived Test 1b matrix contains more than page classifications. Its “Codebook & Sources” sheet explicitly records provenance, assumptions and freeze hashes.
For the primary page/locus metadata source it records:
| Field | Frozen value |
|---|---|
| Source | Zandbergen–Landini |
| Version / item | ZL3b-n, 2025-05-13 |
| Use | Primary page/locus metadata |
| Fields | $Q $P $F $B $I $L $H; locus types |
| Origin URL | https://www.voynich.nu/data/ZL3b-n.txt |
| Expected SHA-256 | bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc |
| Release state | frozen |
The same row carries an independence caveat: this is a single maintained transliteration, while several page classifications contained in the representation are prior scholarly assignments. [2]
That caveat survives this volume unchanged.
A source lock is useful only if it locks both the identity and the limitations. It would be bad version control to preserve the digest while forgetting that the content contains inherited classifications and editorial decisions.
Volume 3 therefore does not rewrite the matrix into a neutral ground truth. It treats the matrix as what its own source sheet says it is: a frozen analytical representation whose provenance can be tested.
3. The First Cross-Check: Exact Digest Agreement
A public Voynich-data project maintained in a separate repository records a source-verification report dated 17 January 2026. In that report, ZL3b-n.txt is listed as accessible, in IVTFF 2.0 with Extended EVA, with a recorded size of 411,671 bytes and the exact SHA-256:
bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc
The same repository’s checksum ledger contains the line:
bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc ZL3b-n.txt
That digest agrees character-for-character with the digest frozen into the eduKate Test 1b matrix. [2–4]
This is stronger than merely finding the same filename on another website. It is a cryptographic identity cross-check between our frozen expected source and a separately preserved verification ledger.
We should still be precise about the word independent. The two records are not independent historical witnesses to the Voynich Manuscript. They are records concerning the same modern ZL3b transliteration. Their agreement does not multiply the evidential weight of the underlying transcription decisions.
What is independent is the preservation route and recordkeeping context. The checksum was not invented by this article to agree with the matrix. It exists in another maintained public research repository.
That is exactly the kind of independence appropriate to a source-identity question.
4. The Second Cross-Check: Header and Provenance Agreement
A separate cipher-benchmark project also preserves a local copy of ZL3b-n.txt. Its source notes identify the same raw source URL on voynich.nu and preserve the source header:
=IVTFF Eva- 2.0 M 5
ZL transliteration file, updated from EVMT project
Version 3b of 13/05/2025
The preserved file itself exposes that header, and the repository documents its importer separately. [5–6]
This gives a second kind of agreement. The name, source URL, format family, version number and version date all align with the source identity frozen in our matrix.
Again, the conclusion must remain bounded. We did not use this second repository as a substitute checksum witness for a digest it does not explicitly state in the source note inspected here. Its value in this volume is provenance and header corroboration.
It also supplies an important methodological warning. Its documentation explains that the benchmark transformation makes choices about alternatives, comments and word boundaries. Those transformation choices belong to the benchmark representation, not automatically to ZL3b itself. [5]
That is a useful reminder for our own next step. Once exact source identity is established, the parser and analytical-view rules become the next source of divergence.
5. What Volume 3 Can Now Mark PASS
The new-generation series uses noncompensatory gates. A strong result in one lane cannot pay for a missing result in another.
Within that framework, Volume 3 can mark the following narrow claims as passed:
- PASS — Frozen expected source identity: the Test 1b matrix explicitly records ZL3b-n, version 3b, dated 13 May 2025, with a full SHA-256.
- PASS — External preserved checksum correspondence: a separately maintained verification record gives the exact same SHA-256 for ZL3b-n.txt.
- PASS — Version/header correspondence: an additional preservation route exposes the same ZL3b version/date header and identifies the same upstream URL.
- PASS — No collision: this volume records source identity and does not compete with Vol. 2’s matrix audit or the older thematic transcription articles.
The first new-generation source-identity receipt can therefore be labelled:
EDKSG-VOY-V003-SOURCE-01
subject: ZL3b-n.txt
frozen_expected_sha256: bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc
preserved_verification_sha256: bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc
disposition: PASS_CROSS_RECORD_SOURCE_IDENTITY
origin_host_current_bytes: HOLD_NOT_REHASHED_IN_THIS_RUN
This is a local research-record identifier. It does not claim to be an identifier issued by the source owner, Yale, a repository, or another eduKate canonical runtime.
6. What Must Remain HOLD
The most important part of this volume may be what the matching digest does not allow us to say.
Current origin-host byte identity
We have not, in this run, directly acquired the complete bytes currently served by the primary voynich.nu origin and calculated their SHA-256 ourselves. The cross-check is against preserved public records and preserved copies. That is enough to validate the frozen source correspondence; it is not enough to claim that the origin host has not changed since those snapshots.
Current-host identity remains HOLD until a complete fresh source capture can be retained and hashed.
Transcription accuracy against the manuscript
A hash verifies which transcription file is under discussion. It cannot determine whether every visible mark was transcribed correctly. That question requires comparison with manuscript images and explicit policies for uncertain forms.
Transcription accuracy remains bounded by the source’s own scholarly status and by any independent image-level audit actually performed.
Character ontology
The cipher-benchmark source notes preserve a warning attributed to the transliteration source: the transliteration alphabet and grouping of glyphs into characters are choices for representation. Even a perfect byte match cannot turn those choices into final historical ontology. [5]
This matters enormously for Voynich. If our analytical units are wrong, a precisely reproducible statistic can be precisely reproducible over the wrong units.
The inherited q/y statistics
Volume 2 recorded earlier reported boundary odds ratios but did not recover the complete event tables and code needed to reproduce those figures. A verified source identity is a prerequisite for reconstruction. It is not the reconstruction itself.
The q/y test remains HOLD_NOT_RERUN.
Test 2A visual-component coding
The source cross-check concerns a textual transliteration representation. It does not complete the separate visual-component experiment. Test 2A remains open.
Semantics, provenance and decipherment
No source hash can identify a word, a language, an author, a city, a function or a plaintext. Those claims require their own evidence gates.
The manuscript remains undeciphered in this volume.
7. The Important Difference Between Identity and Authority
Source identity and source authority are related but different.
Identity asks: which exact representation did we use?
Authority asks: why should this representation be trusted for the claim being made?
For a page/locus metadata analysis, the ZL transliteration is a maintained scholarly resource with explicit structure and provenance. That makes it a serious input. But its adequacy depends on the job.
If the question is whether a page header encodes a particular quire identifier, a structured transcription can be highly useful. If the question is whether two nearly touching strokes were made as one glyph, the public transcription may not carry the physical evidence needed to settle it.
Authority is therefore claim-scoped.
The source can be authoritative for its own maintained transcription state without being authoritative for every interpretation downstream.
This is why the new-generation volumes preserve the route:
object → representation → versioned source → parser → analytical view → statistic → interpretation
A failure anywhere in that route changes what the final statement is allowed to claim.
8. The Hash Does Not Make the Transcription Neutral
There is a subtle trap in reproducibility work. Once an exact file has been identified, the analyst can begin to treat the file as reality.
But the file is still a representation.
A transliteration contains decisions about how visible forms are grouped, how alternatives are represented, how spaces and uncertain spaces are encoded, how comments are stored, how line and paragraph structures are marked, and which metadata are attached to the page.
The Test 1b matrix itself records this problem. Its source caveat says the transliteration is maintained but not independent of prior scholarly assignments. Its non-negotiable limitations state that the distributional clusters are not decoded topics, that stable word-form identity is an assumption rather than an established fact, and that several analytical axes are partly dependent. [2]
Volume 3 does not weaken those cautions because the hash matched.
It strengthens them.
We now know more precisely which representation carries those assumptions.
9. The Parser Becomes the Next Provenance Layer
Once the exact source identity is stable, two analysts can still produce different datasets from it.
One may split uncertain spaces. Another may join them. One may take the first alternative reading. Another may preserve all alternatives. One may remove comments before tokenisation. Another may use comment metadata as a filter. One may treat drawing interruptions as boundaries. Another may bridge them.
The underlying bytes are identical. The analytical populations are not.
This is why the next volume must bind parser behaviour as explicitly as Volume 3 binds source identity.
The cipher-benchmark project is useful as a control precisely because it documents its own transformation choices: it chooses the first alternative in bracketed alternatives, removes IVTFF comments and paragraph/drawing markup for its canonical text, treats both certain and uncertain spaces as word boundaries, and maps units into its own token scheme. [5]
Those choices may be appropriate for that benchmark’s job. They are not automatically our choices.
Our parser needs a different first principle: preserve the source distinction before generating the analytical view. The untouched source layer should remain recoverable so that a later change in separator policy does not require reacquiring the corpus or pretending that the earlier parser saw something it did not.
10. A Three-Layer Source Contract for the Next Run
The next computation should carry three separate identities.
Layer A — Source bytes
This is the exact ZL3b byte sequence used. Its identity is the SHA-256. If we later acquire a fresh current-host source whose digest differs, that becomes a new input edition.
Layer B — Parsed source representation
This retains page identifiers, locus identifiers, source position codes, certain and doubtful separators, alternatives, uncertainty markers and other fields required for the intended test.
Layer C — Analytical view
This applies the declared experimental rules: eligible populations, matched strings, boundary classes, exclusions, group assignments and development/evaluation partitions.
The statistic belongs to Layer C.
The representation choices belong to Layer B and Layer C.
The historical manuscript belongs to neither.
Keeping the layers separate makes later correction possible without erasing history.
11. Why Current-Host Identity Still Matters
The preserved checksum correspondence is enough to validate the identity of the historical source used by the frozen matrix. So why bother checking the current origin later?
Because research sources are living objects too.
A maintained file can be corrected. A server can replace a file while retaining its familiar URL. A header may remain the same while a small textual correction changes the byte sequence. None of those events is necessarily a problem. The problem is pretending they did not happen.
If the current origin digest eventually matches the frozen digest, the result is simple: the currently retrieved bytes agree with the historical source identity used by the matrix.
If it differs, the correct response is not to choose whichever file reproduces the preferred old statistic. The correct response is to retain both source identities, inspect the diff, determine whether the change is editorial or substantive for our test, and label the analytical runs accordingly.
A changed source can be better. Version control is not an argument for preserving errors. It is an argument for making the improvement traceable.
12. Cross-Repository Agreement Is Not Independent Voynich Evidence
There is another subtle counting error to prevent.
Suppose three repositories preserve the same ZL3b transcription. We now have three copies of the source.
We do not have three independent transcriptions of the manuscript.
The copies are extremely useful for preservation and source identity. They can reveal file drift, broken links and accidental substitutions. They cannot be counted as three votes that a particular glyph was read correctly.
For transcription sensitivity, we need genuinely different transcription histories or returns to the manuscript images. Volume 2 already noted that the Takahashi-derived IT source and other transcription families can serve different comparison jobs, provided their ancestry is made explicit.
This principle generalises to the whole Voynich programme:
More copies of one upstream observation increase resilience. They do not automatically increase evidential independence.
That rule protects source preservation from being mistaken for scientific replication.
13. The Exact State Change From Volume 2 to Volume 3
The new series is useful only if each volume creates a visible research-state transition.
| Question | State after Vol. 2 | State after Vol. 3 |
|---|---|---|
| Frozen matrix source filename/version | Identified | Identified |
| Frozen matrix expected SHA-256 | Known in source sheet | Explicitly cross-checked |
| Independent preserved checksum correspondence | Not yet closed | PASS |
| Version/date header correspondence | Candidate inspected | Corroborated across preservation route |
| Current primary-host bytes re-hashed today | Not established | HOLD |
| Original q/y event tables reconstructed | Not completed | HOLD |
| Final-y experiment rerun | Not run | HOLD |
| Test 2A completed | Not run | HOLD |
| Plaintext / semantics / provenance | No new identification | No new identification |
This is a small state transition in the right direction. One source-identity uncertainty has become a recorded correspondence. The larger interpretive uncertainties have not been allowed to ride through the gate behind it.
14. What This Changes About the q/y Test
It changes the test’s starting condition.
Before this cross-check, an attempted reconstruction could fail and leave us unsure whether the disagreement came from the analysis rules or from a different source file.
Now we have a historically frozen expected source identity with an external preserved checksum correspondence. That gives the reconstruction a stable reference target.
The next q/y work can therefore ask a cleaner sequence of questions:
- Are the source bytes used for the reconstruction the frozen ZL3b source identity?
- Does the parser preserve the IVTFF distinctions needed by the experiment?
- Which loci and tokens are eligible?
- How are uncertain gaps and alternatives handled?
- How are current-token and preceding-token endings separated?
- What is the matched-string definition?
- Which page populations are included?
- Which physical or analytical groups define dependence and evaluation splits?
- Can the old event table be reconstructed under the documented rules?
- If not, where does the divergence first appear?
Only after those questions are answered should a numerical effect become the headline.
The source cross-check has not made the q/y hypothesis stronger. It has made the next attempt to test it more interpretable.
15. What This Does Not Change About the Visual Test
Almost nothing.
Test 2A concerns recurring visual components in manuscript images. Its source requirements involve image authority, crop provenance, blind-coding procedure, matched controls and independence conditions.
A verified ZL3b text-source identity is not evidence that a proposed visual component recurs.
The two lanes may eventually meet. If a recurring visual component is independently established and a text feature is independently established, the project can test whether the relationship between them survives confounders and held-out material.
Until then, each lane keeps its own evidence budget.
This is the point of the noncompensatory Wintour gate: success in the text-source lane cannot buy a pass in the visual lane.
16. The Next Research Job
Volume 2 proposed two adjacent tasks: close source identity and validate a location-preserving extraction before measuring the q/y association.
Volume 3 closes the first task within the scope stated above.
The next justified job is therefore parser and locus preservation.
The programme should build or inspect an extraction that preserves enough of the IVTFF source structure to reproduce the intended boundary comparisons without silently redefining the data. At minimum, the record needs source page, locus identity, locus type, paragraph state, exact source text, alternatives or uncertainty state, separator classes, drawing interruptions and the lineage from source text to analytical tokens.
Then the parser should face deterministic tests built around intentionally awkward cases: certain spaces, doubtful spaces, alternative readings, comments, paragraph boundaries, drawing interruptions, empty or nontext loci, special characters and source metadata.
The test suite should verify parser behaviour. It should not be described as scientific validation of a Voynich hypothesis.
If that parser layer passes, the next volume can freeze an analytical view for one bounded boundary test. If it does not pass, the parser failure itself becomes the next research result.
17. The Wintour v1.0 Publication Gate
This volume is intentionally less spectacular than a proposed translation. That is part of its editorial quality.
The current Wintour House publication framework separates commission, evidence, manuscript, edition, proof, release and correction. The new-generation Voynich volumes apply that structure to research state.
For Volume 3, the hard gates are:
| Gate | Disposition |
|---|---|
| Owner / series identity | PASS — Voynich Research Library / EDKSG-VOY series |
| Distinct reader job | PASS — source identity cross-check |
| Predecessor binding | PASS — Vol. 2 next action |
| Frozen source evidence | PASS |
| External preserved checksum correspondence | PASS |
| Current origin bytes | HOLD, explicitly scoped out of PASS |
| Rights / redistribution | PASS for this article — metadata, hashes and short source headers only; no raw transcription corpus republished |
| Semantic promotion | BLOCKED |
| Decipherment promotion | BLOCKED |
| Edition lineage | PASS — V001 → V002 → V003 |
The result is publishable because the article’s claim is smaller than the evidence supporting it.
That is the governing rule of this generation.
18. World Return: We Know Which File We Mean
The Voynich Manuscript remains unread.
Volume 3 has not discovered a word.
It has done something more basic.
It has reduced one source-level ambiguity before the next calculation.
The frozen Test 1b matrix says which ZL3b source it expected. A separately maintained verification record carries the same exact SHA-256. Another public preservation route carries the matching source URL and version header. The source lineage is therefore better anchored than it was at the end of Volume 2.
The result is deliberately bounded. We have not claimed current-origin byte identity without re-hashing the origin. We have not confused a source hash with transcription accuracy. We have not confused transcription accuracy with symbol ontology. We have not confused symbol ontology with language. We have not confused language with meaning.
Those separations are not obstacles around the research.
They are the route through it.
Before asking what the symbols mean, make sure tomorrow’s test uses the same symbols, from the same source, under rules we can reproduce.
The next volume should therefore leave the source-identity gate behind and move one layer downstream: from exact source to exact parser behaviour.
Sources, Evidence Boundaries and Continuation
[1] Predecessor. Voynich | edkSG Vol. 2 — The Input Audit. The source-identity task is inherited from its “next justified step”.
[2] Frozen internal source record. Voynich_Test_1b_Frozen_Segmentation_Matrix.xlsx, “Codebook & Sources”, frozen 2 August 2026. It records Zandbergen–Landini ZL3b-n, 2025-05-13, origin URL https://www.voynich.nu/data/ZL3b-n.txt, and expected SHA-256 bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc. The matrix also preserves source-dependence and representation caveats.
[3] Public verification report. VCAT Source Verification Report, dated 17 January 2026. It records ZL3b-n.txt, version 3b, 411,671 bytes, IVTFF 2.0 with Extended EVA, complete-manuscript coverage and SHA-256 bf5b6d4ac1e3a51b1847a9c388318d609020441ccd56984c901c32b09beccafc.
[4] Public checksum ledger. checksums_all.txt, containing the full ZL3b SHA-256 beside ZL3b-n.txt. This is a preserved verification record, not independent transcription of the manuscript.
[5] Second preservation route and transformation notes. Voynich ZL3b Source Notes. The record identifies the same upstream source URL, a preserved local ZL3b snapshot, the 3b/13-05-2025 source header and explicit benchmark import choices.
[6] Preserved ZL3b file. cipher_benchmark preserved ZL3b-n.txt. Used here for source-header and provenance corroboration; this volume does not claim to have independently calculated its SHA-256 from repository bytes.
Origin-host status. The primary source URL remains the maintained voynich.nu route recorded in the frozen matrix. This edition does not claim a fresh complete-byte download and hash of the current origin-host file. Current origin identity is therefore intentionally held open.
Next: EDKSG-VOY-V004 should address parser and locus preservation: exact source → IVTFF-preserving parsed layer → frozen analytical view, before any q/y statistic is promoted.
Catalogue: Voynich Research Library · Editorial framework: Wintour House.
EDKSG-VOY-V003 · Edition 1.0.0 · 5 September 2026 · Predecessor EDKSG-VOY-V002 · Baseline EDKSG-VOY-B000. Source-identity correspondence PASS within scope. Current primary-host byte identity, q/y reconstruction, Test 2A, semantics, provenance and decipherment remain HOLD or unchanged.