Key points are not available for this paper at this time.
The rapid growth of image-based multimedia content on the Web has intensified the challenge of generating high-quality alternative (alt) text descriptions, which is an essential requirement for inclusive online experiences for people with visual impairments. Although recent advances in machine learning (ML) have enabled large-scale automated alt text generation, the accessibility value of such outputs remains limited. This is due to the context-agnostic datasets used to train existing models, resulting in generic descriptions that fail to meet users’ needs in alt text. In this work, we introduce and utilise a human-curated, context-driven dataset of alt text descriptions to train two proof-of-concept ML models aimed at improving alt text quality. We evaluate these models within a controlled, reproducible pipeline and demonstrate that context-aware training leads to statistically significant improvements in human-perceived alt text quality compared to a model trained without contextual inputs. We further examine the role of context-dependent routing and the integration of contextual cues in shaping generated descriptions, both of which are critical but underexplored aspects of alt text accessibility. The findings highlight the value of structured, human-curated contextual data in advancing ML-supported alt text generation and point towards opportunities for hybrid human-AI approaches to inclusive web design.
Droutsas et al. (Fri,) studied this question.