This computational analysis reveals the constrained morphological structure of the Torah, indicating its unique linguistic properties.
This study presents a computational morphological analysis of the Torah (Five Books of Moses) as a closed linguistic corpus of 76,584 tokens. Every word is decomposed into three structural layers: a foundation root (semantic anchor), a mandatory root (stable consonantal skeleton), and an expansion layer (morphological extensions). The central finding: morphological expansion is governed by a restricted "control alphabet" of 10 letters -- EmetNiyahu (aleph, mem, tav, nun, yod, he, vav) plus BKL (bet, kaf, lamed) -- accounting for 99.87% of all 97,599 extension tokens (p <= 0.0003, 0/10,000 random sets). A 30-line extraction algorithm, trained on 80% of the Torah with no external dictionary (data: Sefaria.org API only), predicts the semantic group of unseen words with 90.1% accuracy (5-fold CV, sigma=0.2%). Adding nikud (vowel pointing) raises accuracy to 95.6%, quantifying the exact information content preserved by the oral tradition. The AMTN letters form an independent parallel root system with 57.6% compositional decomposition and 96.6% YHW-based meaning separation. A shuffle test (Z=57.72, 0/1,000) confirms that foundation-letter clustering is a property of the specific text, not of Hebrew in general. Cross-biblical comparison reveals a clear hierarchy (Torah Z=25-31 > Prophets > Writings), while Biblical Aramaic shows the same YHW mechanism but no significant narrative clustering (Z=0.39). All data from Sefaria.org. Complete code provided. Fully reproducible.
No takes yet. Share an insight, caveat, or question.
ERAN ELIYAHU Tobul (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: