Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Yong, Li, Chenxing, Xu, Le, Gu, Hao, Zhang, Duzhen, Chen, Yujie, Xu, Manjie, Fu, Ruibo, Yang, Shan, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918036724776960
author Ren, Yong
Li, Chenxing
Xu, Le
Gu, Hao
Zhang, Duzhen
Chen, Yujie
Xu, Manjie
Fu, Ruibo
Yang, Shan
Yu, Dong
author_facet Ren, Yong
Li, Chenxing
Xu, Le
Gu, Hao
Zhang, Duzhen
Chen, Yujie
Xu, Manjie
Fu, Ruibo
Yang, Shan
Yu, Dong
contents Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Ren, Yong
Li, Chenxing
Xu, Le
Gu, Hao
Zhang, Duzhen
Chen, Yujie
Xu, Manjie
Fu, Ruibo
Yang, Shan
Yu, Dong
Multimedia
Sound
Audio and Speech Processing
Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference.
title Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.13062