Randomized trial analyzes information metrics in the Voynich Manuscript, suggesting characteristics of its structure and language evolution.
Key Points
This paper aims to analyze the Voynich Manuscript's linguistic characteristics using various information-theoretic measurements.
Applied a layered measurement battery to the Zandbergen-Landini ZL3b transliteration (34,103 running-text tokens) and compared it against seven medieval reference corpora.
Quantified noise floors for character entropy and word-order analysis to evaluate linguistic properties.
Simulated Latin scribal abbreviation effects and developed a generative model based on the manuscript's text.
Character entropy for the manuscript is significantly lower at h2 = 2.17 bits compared to medieval reference corpora (3.02-3.52 bits).
Word reuse decreases across folio and paragraph boundaries, indicating that repetition depends on the physical page layout rather than the linear text.
The generative model reproduces 9 of 14 target statistics, indicating systematic residuals suggesting stronger real bindings compared to those the model predicts.