Randomized trial compares OCR reliability in noisy documents, suggesting optimized model architecture improves outcomes.
Optical character recognition (OCR) is an essential technique for transforming document images into readable text. However, the OCR process faces significant challenges when processing moderately noisy, unstructured documents due to practical degradations such as blurring, stain marks, ink degradation, and background noise. In this research, a comprehensive comparative analysis of three state-of-the-art deep learning models, namely EasyOCR, Donut, and LayoutLMv3, was conducted on the Kaggle Denoising Dirty Documents dataset before implementing a hybrid combiner based on mutual similarity voting and three post-processing algorithms consisting of rule-based correction, Retrieval-Augmented Generation (RAG), and Large Language Model-based correction. The experimental analysis was conducted by evaluating all models on 20 test images using Word Error Rate (WER), Character Error Rate, Accuracy, Precision, Recall, F1-score, and Confusion Matrix as performance metrics. The results show that LayoutLMv3 achieved the best F1-score of 0.9168 by integrating text, image patches, and spatial coordinates within a transformer network. EasyOCR attained the lowest WER of 0.2883 using a two-stage approach consisting of CRAFT detection and Convolutional Recurrent Neural Network-based text recognition. Donut also performed poorly on the dataset, achieving a WER of 2.0997, showing that pretraining end-to-end models on structured documents fails when applied to unstructured, noisy paragraphs - an important negative insight regarding model-task alignment. The hybrid combiner model produced consistently strong performance, achieving a WER of 0.3192 and an F1-score of 0.8995 by effectively eliminating inaccurate Donut predictions through a similarity-based voting scheme. In terms of the post-processing methods, rule-based correction achieved the greatest improvement in WER, reaching 0.3185, while the use of RAG led to degraded performance, with a WER of 0.5630, owing to a small corpus size that resulted in incorrect document retrieval. These findings demonstrate that selecting an architecturally appropriate model - specifically one that jointly encodes text, image, and spatial layout - delivers greater performance gains than applying complex post-processing, providing clear design guidance for deploying OCR systems on moderately noisy, unstructured documents.
No takes yet. Share an insight, caveat, or question.
Trivedi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: