Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karamolegkou, Antonia, Nikandrou, Malvina, Pantazopoulos, Georgios, Villegas, Danae Sanchez, Rust, Phillip, Dhar, Ruchira, Hershcovich, Daniel, Søgaard, Anders
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909556291928064
author Karamolegkou, Antonia
Nikandrou, Malvina
Pantazopoulos, Georgios
Villegas, Danae Sanchez
Rust, Phillip
Dhar, Ruchira
Hershcovich, Daniel
Søgaard, Anders
author_facet Karamolegkou, Antonia
Nikandrou, Malvina
Pantazopoulos, Georgios
Villegas, Danae Sanchez
Rust, Phillip
Dhar, Ruchira
Hershcovich, Daniel
Søgaard, Anders
contents This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, our findings highlight concerns related to contextual understanding, cultural sensitivity, and complex scene understanding, particularly for individuals who may rely solely on them for visual interpretation. Informed by these results, we collate five user-centred tasks with image and video inputs, including a novel task on Optical Braille Recognition. Our systematic evaluation of twelve MLLMs reveals that further advancements are necessary to overcome limitations related to cultural context, multilingual support, Braille reading comprehension, assistive object recognition, and hallucinations. This work provides critical insights into the future direction of multimodal AI for accessibility, underscoring the need for more inclusive, robust, and trustworthy visual assistance technologies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
Karamolegkou, Antonia
Nikandrou, Malvina
Pantazopoulos, Georgios
Villegas, Danae Sanchez
Rust, Phillip
Dhar, Ruchira
Hershcovich, Daniel
Søgaard, Anders
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, our findings highlight concerns related to contextual understanding, cultural sensitivity, and complex scene understanding, particularly for individuals who may rely solely on them for visual interpretation. Informed by these results, we collate five user-centred tasks with image and video inputs, including a novel task on Optical Braille Recognition. Our systematic evaluation of twelve MLLMs reveals that further advancements are necessary to overcome limitations related to cultural context, multilingual support, Braille reading comprehension, assistive object recognition, and hallucinations. This work provides critical insights into the future direction of multimodal AI for accessibility, underscoring the need for more inclusive, robust, and trustworthy visual assistance technologies.
title Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2503.22610