Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Junyu, Ma, Ziyang, Luo, Zhengding, Wang, Tianrui, Ge, Meng, Wang, Xiaobao, Wang, Longbiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914052235591680
author Wang, Junyu
Ma, Ziyang
Luo, Zhengding
Wang, Tianrui
Ge, Meng
Wang, Xiaobao
Wang, Longbiao
author_facet Wang, Junyu
Ma, Ziyang
Luo, Zhengding
Wang, Tianrui
Ge, Meng
Wang, Xiaobao
Wang, Longbiao
contents Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their ability to fully utilize acoustic cues, causing suboptimal performance on audio reasoning tasks. To mitigate this, we propose \textbf{MATA}, a novel training-free method that dynamically pushes LALMs to pay \textbf{M}ore \textbf{A}ttention \textbf{T}o \textbf{A}udio tokens within the self-attention mechanism. Specifically, MATA intervenes post raw attention scoring, targeting only the last token in intermediate layers without introducing additional parameters or computational overhead. Experiments on the MMAU and MMAR benchmarks confirm MATA's effectiveness, with consistent performance gains. Notably, on MMAR, MATA enables an open-source model to surpass the proprietary Gemini 2.0 Flash for the first time. Our work provides an efficient solution to mitigate attention bias and opens a new research direction for enhancing the audio-processing capabilities of multi-modal models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
Wang, Junyu
Ma, Ziyang
Luo, Zhengding
Wang, Tianrui
Ge, Meng
Wang, Xiaobao
Wang, Longbiao
Sound
Computation and Language
Multimedia
Audio and Speech Processing
Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their ability to fully utilize acoustic cues, causing suboptimal performance on audio reasoning tasks. To mitigate this, we propose \textbf{MATA}, a novel training-free method that dynamically pushes LALMs to pay \textbf{M}ore \textbf{A}ttention \textbf{T}o \textbf{A}udio tokens within the self-attention mechanism. Specifically, MATA intervenes post raw attention scoring, targeting only the last token in intermediate layers without introducing additional parameters or computational overhead. Experiments on the MMAU and MMAR benchmarks confirm MATA's effectiveness, with consistent performance gains. Notably, on MMAR, MATA enables an open-source model to surpass the proprietary Gemini 2.0 Flash for the first time. Our work provides an efficient solution to mitigate attention bias and opens a new research direction for enhancing the audio-processing capabilities of multi-modal models.
title Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
topic Sound
Computation and Language
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2509.18816