Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheema, Maryam, Seifi, Hasti, Fazli, Pooyan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908381737910272
author Cheema, Maryam
Seifi, Hasti
Fazli, Pooyan
author_facet Cheema, Maryam
Seifi, Hasti
Fazli, Pooyan
contents Audio descriptions (AD) make videos accessible for blind and low vision (BLV) users by describing visual elements that cannot be understood from the main audio track. AD created by professionals or novice describers is time-consuming and offers little customization or control to BLV viewers on description length and content and when they receive it. To address this gap, we explore user-driven AI-generated descriptions, enabling BLV viewers to control both the timing and level of detail of the descriptions they receive. In a study, 20 BLV participants activated audio descriptions for seven different video genres with two levels of detail: concise and detailed. Our findings reveal differences in the preferred frequency and level of detail of ADs for different videos, participants' sense of control with this style of AD delivery, and its limitations. We discuss the implications of these findings for the development of future AD tools for BLV users.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11835
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
Cheema, Maryam
Seifi, Hasti
Fazli, Pooyan
Human-Computer Interaction
Audio descriptions (AD) make videos accessible for blind and low vision (BLV) users by describing visual elements that cannot be understood from the main audio track. AD created by professionals or novice describers is time-consuming and offers little customization or control to BLV viewers on description length and content and when they receive it. To address this gap, we explore user-driven AI-generated descriptions, enabling BLV viewers to control both the timing and level of detail of the descriptions they receive. In a study, 20 BLV participants activated audio descriptions for seven different video genres with two levels of detail: concise and detailed. Our findings reveal differences in the preferred frequency and level of detail of ADs for different videos, participants' sense of control with this style of AD delivery, and its limitations. We discuss the implications of these findings for the development of future AD tools for BLV users.
title Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
topic Human-Computer Interaction
url https://arxiv.org/abs/2411.11835