Randomized trial demonstrates enhanced audio description for visually impaired users, suggesting improved accessibility and engagement.
In recent years, automatic audio description (AD) generation has become an important research domain within accessibility and assistive technology, driven by its potential to enhance content understanding, social integration, and cognitive engagement for individuals with visual impairments (VI). In this paper, we introduce EARS4SEE, a novel multimodal framework for AD generation that integrates semantic video analysis, character tracking, and adaptive temporal segmentation to enhance contextual coherence and narrative fluency. The proposed system integrates multi-stream fusion strategy, leveraging visual, textual, and audio modalities for character-centric, semantically enriched AD. Textual descriptions are synthesized into natural-sounding speech using state-of-the-art text-to-speech (TTS) techniques for an immersive experience. A core contribution of the proposed methodology involves the tracking-based character recognition module, which ensures temporally consistent character identification using an adaptive temporal attention mechanism. The approach mitigates inconsistencies from motion blur, occlusions, and scale variations, improving referential continuity. Additionally, EARS4SEE introduces an automated multimodal video segmentation pipeline, capturing long-range temporal dependencies to improve scene boundary detection and contextual alignment. The experimental evaluation carried out on the MAD-Eval-Named and TV-AD datasets validates the effectiveness of the proposed methodology, which leads to average CIDEr and LLM-AD-eval scores of 24.1 and 3.02, respectively. In addition, when compared to state-of-the-art techniques, the proposed architecture shows superior performances in terms of the CIDEr, with gains in accuracy ranging in the [1.72%, 10.2%] interval and an 8% increase in LLM-AD-eval scores. • Scene-aware, training-free framework for long-form audio description. • Multimodal temporal segmentation captures context across multiple shots. • Adaptive temporal attention improves character identity consistency. • Memory-guided prompting improves narrative coherence across shots. • EARS4SEE outperforms recent training-free AD baselines on two benchmarks.
No takes yet. Share an insight, caveat, or question.
Tapu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: