PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 8, 2026ACM Transactions on Multimedia Computing Communications and Applications2 citations

SEADUNet: A Multilingual Ancient Document Image Binarization using EMCAM Attention Mechanism and SCP

View Full Paper
HGHai GuoLTLingling TongJZJingying Zhao

Key Points

  • This research focuses on enhancing the binarization of multi-script ancient document images using a novel approach.
  • Developed SEADUNet combining multi-scale convolutional attention and spatial-channel reconstruction techniques.
  • Utilized the Multilingual Ancient Document Image Binarization Dataset (MADIBD2024-16) for training and evaluation.
  • Conducted experiments comparing the binarization method against traditional and modern techniques.
  • Achieved an F-Measure of 95.54% and a pseudo F-Measure of 95.98%.
  • Obtained a Peak Signal to Noise Ratio of 20.67 dB.
  • Demonstrated superior performance in handling multi-script binarization challenges.

Abstract

As invaluable resources for historical and cultural studies, ancient manuscripts demand immediate digitization and conservation measures to counteract degradation threats such as paper aging, ink fading, and physical damage. Optical character recognition (OCR) is an important protection method for the digitization of ancient manuscripts, and noise reduction and binarization of ancient manuscripts have significant impacts on their recognition accuracy. The binarization of multi-script ancient document images is confronted with a multitude of challenges, including the diversity of preservation media, improper storage practices, variations in writing styles across different languages, and the intricacies of noise. To tackle these complexities, this paper introduces a novel binarization approach named SEADUNet, which seamlessly combines a multi-scale convolutional attention feature fusion module (EMCAM) with spatial-channel reconstructed convolution techniques. This approach harnesses the power of distinctive multi-scale deep convolutional blocks to markedly enhance feature mapping through multi-scale convolution, while also concentrating on the prominent areas within the images. The EMCAM module, through the use of group convolution and deep convolution, exhibits high efficiency and commendable scalability. Advancing research in this field, the Multilingual Ancient Document Image Binarization Dataset (MADIBD2024-16) offers a rigorously curated collection of 3,200 annotated image pairs spanning 16 distinct historical scripts, The ratio of training set to test set is 8:2, providing a standardized benchmark for evaluating document binarization algorithms. The experimental outcomes are impressive, with the method achieving an F-Measure (FM) of 95.54%, a pseudo F-Measure (p-FM) of 95.98%, a Peak Signal to Noise Ratio (PSNR) of 20.67 dB, and a Distance Reciprocal Distortion (DRD) of 2.59 on a newly established dataset. When compared to both traditional and cutting-edge methods, this architecture proves to be particularly adept at handling the binarization of multi-script ancient document images. Moreover, additional validation on binarization datasets from other ancient scripts has substantiated the universality and practicality of this architecture.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Guo et al. (2026) studied this question.

synapsesocial.com/papers/69acc57d32b0ef16a404fa1fhttps://doi.org/10.1145/3800946
Ask AI
Helpful
Bookmark
Share
View Full Paper