Listening without Looking: Modality Bias in Audio-Visual Captioning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ishikawa, Yuchi, Manabe, Toranosuke, Komatsu, Tatsuya, Aoki, Yoshimitsu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909873598365696
author Ishikawa, Yuchi
Manabe, Toranosuke
Komatsu, Tatsuya
Aoki, Yoshimitsu
author_facet Ishikawa, Yuchi
Manabe, Toranosuke
Komatsu, Tatsuya
Aoki, Yoshimitsu
contents Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two modalities are complementary in current audio-visual captioning models and how robust these models are when one modality is degraded. We address these questions by conducting systematic modality robustness tests on LAVCap, a state-of-the-art audio-visual captioning model, in which we selectively suppress or corrupt the audio or visual streams to quantify sensitivity and complementarity. The analysis reveals a pronounced bias toward the audio stream in LAVCap. To evaluate how balanced audio-visual captioning models are in their use of both modalities, we augment AudioCaps with textual annotations that jointly describe the audio and visual streams, yielding the AudioVisualCaps dataset. In our experiments, we report LAVCap baseline results on AudioVisualCaps. We also evaluate the model under modality robustness tests on AudioVisualCaps and the results indicate that LAVCap trained on AudioVisualCaps exhibits less modality bias than when trained on AudioCaps.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listening without Looking: Modality Bias in Audio-Visual Captioning
Ishikawa, Yuchi
Manabe, Toranosuke
Komatsu, Tatsuya
Aoki, Yoshimitsu
Audio and Speech Processing
Computer Vision and Pattern Recognition
Image and Video Processing
Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two modalities are complementary in current audio-visual captioning models and how robust these models are when one modality is degraded. We address these questions by conducting systematic modality robustness tests on LAVCap, a state-of-the-art audio-visual captioning model, in which we selectively suppress or corrupt the audio or visual streams to quantify sensitivity and complementarity. The analysis reveals a pronounced bias toward the audio stream in LAVCap. To evaluate how balanced audio-visual captioning models are in their use of both modalities, we augment AudioCaps with textual annotations that jointly describe the audio and visual streams, yielding the AudioVisualCaps dataset. In our experiments, we report LAVCap baseline results on AudioVisualCaps. We also evaluate the model under modality robustness tests on AudioVisualCaps and the results indicate that LAVCap trained on AudioVisualCaps exhibits less modality bias than when trained on AudioCaps.
title Listening without Looking: Modality Bias in Audio-Visual Captioning
topic Audio and Speech Processing
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2510.24024