Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Olmos, Adriana, Sinha, Anoop K., Santos, Renelito Delos, Rodriguez, Ruben Rodriguez, Landay, James A., Sepah, Sam S., Nelson, Philip, Kane, Shaun K.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911422172102656
author Olmos, Adriana
Sinha, Anoop K.
Santos, Renelito Delos
Rodriguez, Ruben Rodriguez
Landay, James A.
Sepah, Sam S.
Nelson, Philip
Kane, Shaun K.
author_facet Olmos, Adriana
Sinha, Anoop K.
Santos, Renelito Delos
Rodriguez, Ruben Rodriguez
Landay, James A.
Sepah, Sam S.
Nelson, Philip
Kane, Shaun K.
contents Video content remains largely inaccessible to blind and low-vision (BLV) users. To address this, we introduce a prototype that leverages a multimodal agent - powered by a novel conversational architecture using a multimodal large language model (MLLM) - to provide BLV users with an interactive, accessible video experience. This Multimodal Agent Video Player (MAVP) demonstrates that an interactive accessibility mode can be added to a video through multilayered prompt orchestration. We describe a user-centered design process involving 18 sessions with BLV users that showed that BLV users do not just want accessibility features, but desire independence and personal agency over the viewing experience. We conducted a qualitative study with an additional 8 BLV participants; in this, we saw that the MAVP's conversational dialogue offers BLV users a sense of personal agency, fostering collaboration and trust. Even in the case of hallucinations, it is meta-conversational dialogues about AI's limitations that can repair trust.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04104
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player
Olmos, Adriana
Sinha, Anoop K.
Santos, Renelito Delos
Rodriguez, Ruben Rodriguez
Landay, James A.
Sepah, Sam S.
Nelson, Philip
Kane, Shaun K.
Human-Computer Interaction
Video content remains largely inaccessible to blind and low-vision (BLV) users. To address this, we introduce a prototype that leverages a multimodal agent - powered by a novel conversational architecture using a multimodal large language model (MLLM) - to provide BLV users with an interactive, accessible video experience. This Multimodal Agent Video Player (MAVP) demonstrates that an interactive accessibility mode can be added to a video through multilayered prompt orchestration. We describe a user-centered design process involving 18 sessions with BLV users that showed that BLV users do not just want accessibility features, but desire independence and personal agency over the viewing experience. We conducted a qualitative study with an additional 8 BLV participants; in this, we saw that the MAVP's conversational dialogue offers BLV users a sense of personal agency, fostering collaboration and trust. Even in the case of hallucinations, it is meta-conversational dialogues about AI's limitations that can repair trust.
title Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player
topic Human-Computer Interaction
url https://arxiv.org/abs/2602.04104