PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 11, 2025Scientific Reports3 citationsOpen Access

VI-OCR: “Visually Impaired” optical character recognition pipeline for text accessibility assessment

View Full Paper
QGQingying GaoRMRoberto ManduchiPRPradeep Y. Ramulu

Key Points

  • This research aims to assess text accessibility for individuals with low vision using a new optical character recognition pipeline.
  • Developed the VI-OCR pipeline based on advanced OCR models.
  • Benchmark experiments were conducted across three tasks: letter acuity, word acuity, and scene text recognition.
  • Compared the performance of OCR models to that of normal vision participants.
  • Identified major limitations in existing OCR models, such as generalizability and performance under contrast reduction.
  • Highlighted robust human-like performance in advanced models like Qwen2.5-VL and GPT.

Abstract

Low vision adversely impacts daily activities, particularly reading. However, quantifying text accessibility for different levels of low vision is challenging, leading to product designs that often overlook the vision status of low vision readers. In this paper, we bridge the gap between computer vision and low vision fields by introducing a text accessibility assessment pipeline called VI-OCR (short for Visually Impaired Optical Character Recognition), based on state-of-the-art OCR models. VI-OCR mimics human text recognition ability under specified levels of visual acuity and contrast sensitivity loss, to estimate whether text of a given size would be recognizable for a low vision human reader. We benchmarked specialized OCR models and vision-language models in replicating text recognition performances with visual acuity and contrast sensitivity deficits across three reading tasks: letter acuity using ETDRS charts, word acuity using MNREAD charts, and scene text recognition using complex real-life images. Comparing model performance to that of normal vision participants on degraded texts revealed major issues in some models including limited generalizability across reading tasks, difficulties dealing with severe contrast reduction, and overperforming rather than mimicking human observers. However, robust human-like performance of winning models such as Qwen2.5-VL and GPT supports the feasibility of VI-OCR in assessing text accessibility.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gao et al. (2025) studied this question.

synapsesocial.com/papers/6940192a2d562116f28f6b93https://doi.org/10.1038/s41598-025-30982-7
Ask AI
Helpful
Bookmark
Share
View Full Paper