Key points are not available for this paper at this time.
For people with visual impairments, navigating and orienting themselves in public spaces is a daily challenge. Although a wide variety of electronic assistive devices already exist, in the past these provided only fairly abstract guidance. With large language models (LLMs), natural-language descriptions of image content are now possible. While there are already approaches that capture images via smart glasses and forward the image content to external large language models to generate image descriptions, these methods raise concerns regarding data privacy and availability. This paper therefore presents a novel approach in which data is processed exclusively locally on the user’s device. To do this, the user captures a camera image via smart glasses, which is forwarded to a connected smartphone. Inference takes place directly on the smartphone using a reduced, local LLM that has been adapted through LoRA fine-tuning using data relevant to people with visual impairments. The generated image description is then played through a speaker in the temple of the smart glasses. For the task-specific fine-tuning, a comprehensive survey of visually impaired people was conducted, and professional mobility trainers were consulted. The study demonstrated the feasibility of fully local processing using smart glasses in combination with a smartphone. The evaluation showed that fine-tuning the small, local model yields improved image descriptions compared to the non-adapted model, some of which outperform a significantly larger external language model.
Götzelmann et al. (Wed,) studied this question.