More than a Moment: Towards Coherent Sequences of Audio Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khandelwal, Eshika, Xie, Junyu, Han, Tengda, Bain, Max, Nagrani, Arsha, Zisserman, Andrew, Varol, Gül, Tapaswi, Makarand
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911239489191936
author Khandelwal, Eshika
Xie, Junyu
Han, Tengda
Bain, Max
Nagrani, Arsha
Zisserman, Andrew
Varol, Gül
Tapaswi, Makarand
author_facet Khandelwal, Eshika
Xie, Junyu
Han, Tengda
Bain, Max
Nagrani, Arsha
Zisserman, Andrew
Varol, Gül
Tapaswi, Makarand
contents Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that helps listeners to visualise the unfolding scene, rather than describing isolated moments. However, most automatic methods generate each AD independently, often resulting in repetitive, incoherent descriptions. To address this, we propose a training-free method, CoherentAD, that first generates multiple candidate descriptions for each AD time interval, and then performs auto-regressive selection across the sequence to form a coherent and informative narrative. To evaluate AD sequences holistically, we introduce a sequence-level metric, StoryRecall, which measures how well the predicted ADs convey the ground truth narrative, alongside repetition metrics that capture the redundancy across consecutive AD outputs. Our method produces coherent AD sequences with enhanced narrative understanding, outperforming prior approaches that rely on independent generations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25440
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle More than a Moment: Towards Coherent Sequences of Audio Descriptions
Khandelwal, Eshika
Xie, Junyu
Han, Tengda
Bain, Max
Nagrani, Arsha
Zisserman, Andrew
Varol, Gül
Tapaswi, Makarand
Computer Vision and Pattern Recognition
Computation and Language
Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that helps listeners to visualise the unfolding scene, rather than describing isolated moments. However, most automatic methods generate each AD independently, often resulting in repetitive, incoherent descriptions. To address this, we propose a training-free method, CoherentAD, that first generates multiple candidate descriptions for each AD time interval, and then performs auto-regressive selection across the sequence to form a coherent and informative narrative. To evaluate AD sequences holistically, we introduce a sequence-level metric, StoryRecall, which measures how well the predicted ADs convey the ground truth narrative, alongside repetition metrics that capture the redundancy across consecutive AD outputs. Our method produces coherent AD sequences with enhanced narrative understanding, outperforming prior approaches that rely on independent generations.
title More than a Moment: Towards Coherent Sequences of Audio Descriptions
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.25440