PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 2026Machine Learning and Knowledge Extraction0 citationsOpen Access

Document Image Binarization Using Various Machine Learning Models and Ensembles Trained on Classic Local and Global Binarization Algorithms and Image Statistics

View Full Paper
NTNicolae TarbăCBCostin-Anton BoiangiuMVMihai-Lucian Voncilă

Key Points

  • This research aims to improve document image binarization by combining machine learning with traditional thresholding methods to effectively handle noise.
  • Developed a mixed global-local thresholding method leveraging machine learning frameworks.
  • Utilized results from various binarization algorithms and image statistics for training.
  • Conducted cross-validation to evaluate robustness and performance on new datasets.
  • Achieved results comparable to advanced methods on benchmark document image binarization datasets.

Abstract

Image binarization is a preprocessing technique that maps an image’s pixel values to either black or white, and it is crucial in many fields of computer vision, such as document digitization and medical imaging. Thresholding is a popular image binarization technique for grayscale images because it splits pixel values into greater than or lower than a specific threshold. Global thresholding is fast because it computes only one threshold for the entire image, but it cannot handle many types of noise specific to document images. Local thresholding has greater computational complexity because it adjusts the thresholds for each pixel based on the surrounding pixels, but it can handle such types of noise, although it risks introducing noise in uniform areas of the image. Mixed global–local approaches can mitigate this risk while still being able to handle most types of noise. This paper proposes a mixed global–local thresholding method that harnesses two popular automatic machine learning frameworks to train machine learning models using the results of several thresholding algorithms and other image statistics. Cross-validation was performed to ensure that the selected models are robust and perform well on new data. We obtained results comparable with other state-of-the-art methods on popular document image binarization datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tarbă et al. (2026) studied this question.

synapsesocial.com/papers/6a1fc42cdee9eb8c0dce5b54https://doi.org/10.3390/make8060149
Ask AI
Helpful
Bookmark
Share
View Full Paper