When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Cheng, Deng, Gelei, Yang, Xianglin, Qiu, Han, Zhang, Tianwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908497110630400
author Wang, Cheng
Deng, Gelei
Yang, Xianglin
Qiu, Han
Zhang, Tianwei
author_facet Wang, Cheng
Deng, Gelei
Yang, Xianglin
Qiu, Han
Zhang, Tianwei
contents Large Audio-Language Models (LALMs) are enhanced with audio perception capabilities, enabling them to effectively process and understand multimodal inputs that combine audio and text. However, their performance in handling conflicting information between audio and text modalities remains largely unexamined. This paper introduces MCR-BENCH, the first comprehensive benchmark specifically designed to evaluate how LALMs prioritize information when presented with inconsistent audio-text pairs. Through extensive evaluation across diverse audio understanding tasks, we reveal a concerning phenomenon: when inconsistencies exist between modalities, LALMs display a significant bias toward textual input, frequently disregarding audio evidence. This tendency leads to substantial performance degradation in audio-centric tasks and raises important reliability concerns for real-world applications. We further investigate the influencing factors of text bias, and explore mitigation strategies through supervised finetuning, and analyze model confidence patterns that reveal persistent overconfidence even with contradictory inputs. These findings underscore the need for improved modality balance during training and more sophisticated fusion mechanisms to enhance the robustness when handling conflicting multi-modal inputs. The project is available at https://github.com/WangCheng0116/MCR-BENCH.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15407
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
Wang, Cheng
Deng, Gelei
Yang, Xianglin
Qiu, Han
Zhang, Tianwei
Computation and Language
Artificial Intelligence
Large Audio-Language Models (LALMs) are enhanced with audio perception capabilities, enabling them to effectively process and understand multimodal inputs that combine audio and text. However, their performance in handling conflicting information between audio and text modalities remains largely unexamined. This paper introduces MCR-BENCH, the first comprehensive benchmark specifically designed to evaluate how LALMs prioritize information when presented with inconsistent audio-text pairs. Through extensive evaluation across diverse audio understanding tasks, we reveal a concerning phenomenon: when inconsistencies exist between modalities, LALMs display a significant bias toward textual input, frequently disregarding audio evidence. This tendency leads to substantial performance degradation in audio-centric tasks and raises important reliability concerns for real-world applications. We further investigate the influencing factors of text bias, and explore mitigation strategies through supervised finetuning, and analyze model confidence patterns that reveal persistent overconfidence even with contradictory inputs. These findings underscore the need for improved modality balance during training and more sophisticated fusion mechanisms to enhance the robustness when handling conflicting multi-modal inputs. The project is available at https://github.com/WangCheng0116/MCR-BENCH.
title When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.15407