Optical character recognition (OCR) technology aims mainly to transform printed or handwritten text into machine-readable text. Recently, OCR technology has improved text scanning for various languages. It is widely used for digitizing physical documents, making them searchable and editable digitally. It aids in language learning and boosts business efficiency. However, attempts to apply OCR technology to Arabic texts face several challenges. This is because of the high complexity of its script, the cursive nature of Arabic characters, and contextual variations. Many Arabic OCR mechanisms have been developed in the literature. However, they suffer high error rates and issues of low accuracy. Then, an accurate OCR benefits the visually impaired, making content accessible via text-to-speech and braille displays is required. This work introduces a new dataset, MFSRHRD (Multiple Fonts, Sizes, Resources, High Records Dataset), containing sufficient records to ensure adequate training and correct learning of the model. This dataset gathers two types of words; the first includes words with diacritics, and the second includes words without diacritics. Besides, the new dataset contains images of the words collected from different specialized websites to ensure the diversity of the words and separate character images with and without diacritics. Therefore, the proposed dataset is classified at word and character levels. All these words and letters were drawn in different font styles, sizes, and image quality. The diacritical marks are dhamma, fatha, kasra, shade, and sukun. This work describes the detailed specifications of the proposed dataset intended for the research community.
No takes yet. Share an insight, caveat, or question.
Gaashan et al. (2024) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: