This lightning talk will present a reproducible workflow for developing a Handwritten Text Recognition (HTR) model tailored to the manuscripts of the Belgian historian Henri Pirenne. Combining generic models with domain-specific fine-tuning, we construct a high-quality ground truth corpus of 473 manually corrected pages spanning different periods and genres of the historian’s work. The workflow integrates transcription, segmentation, and iterative model training using open-source tools, with particular attention to standardized transcription conventions and interoperable data structures. The resulting models achieve over 93% character accuracy, demonstrating the effectiveness of specialization for complex handwriting. This work contributes to transparent HTR practices and enables large-scale digital analysis of historical archives.
Carmen Carrasco Luján (Mon,) studied this question.