This paper raises a methodological question: How can we assess machine-made Old English when there is no parallel reference text and the standard metrics do not fit the task? We propose a pipeline with five measuring layers plus two compliance components, including lexical attestation with form linking, frequency-profile diagnostics, character-level comparison, word-embedding geometry under a verified mapping and dependency parsing, with bootstrap confidence intervals around the main contrasts. We apply the pipeline to a machine-made version of Gregory’s Dialogues: 4217 sentences, one for each sentence of the Old English original, generated under hard constraints. The unattested residue is two word types and 0.003% of tokens. The frequency profile diverges from the original by 0.008, less than the original diverges from the background corpus. At character level, in embedding space and in parsed syntax, the generated text stands at the same distance from the Dictionary of Old English Corpus as the original itself does. We propose an overall metric G, the geometric mean of seven bounded components, which scores the text at 0.991 with a confidence interval of [0.991, 0.992]. Two blind detection experiments with expert judges place the index externally: roughly two-thirds of generated sentences pass as authentic to specialists, so the divergence the pipeline measures is real at corpus scale but not available to sentence-by-sentence reading. The main contribution is a reusable evaluation method for historical language generation, together with a single interpretable score that subsumes the partial metrics without hiding them.
No takes yet. Share an insight, caveat, or question.
Arista et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: