Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918036724776960 |
|---|---|
| author | Ren, Yong Li, Chenxing Xu, Le Gu, Hao Zhang, Duzhen Chen, Yujie Xu, Manjie Fu, Ruibo Yang, Shan Yu, Dong |
| author_facet | Ren, Yong Li, Chenxing Xu, Le Gu, Hao Zhang, Duzhen Chen, Yujie Xu, Manjie Fu, Ruibo Yang, Shan Yu, Dong |
| contents | Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13062 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Ren, Yong Li, Chenxing Xu, Le Gu, Hao Zhang, Duzhen Chen, Yujie Xu, Manjie Fu, Ruibo Yang, Shan Yu, Dong Multimedia Sound Audio and Speech Processing Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference. |
| title | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model |
| topic | Multimedia Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.13062 |