MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Chih-Kai, Tsai, Yun-Shao, Guo, Yu-Kai, Tsai, Ping-Le, Piao, Yen-Ting, Chen, Hung-Wei, Hsiao, Ting-Lin, Hsu, Yun-Man, Lu, Ke-Han, Lee, Hung-yi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917329865015296
author Yang, Chih-Kai
Tsai, Yun-Shao
Guo, Yu-Kai
Tsai, Ping-Le
Piao, Yen-Ting
Chen, Hung-Wei
Hsiao, Ting-Lin
Hsu, Yun-Man
Lu, Ke-Han
Lee, Hung-yi
author_facet Yang, Chih-Kai
Tsai, Yun-Shao
Guo, Yu-Kai
Tsai, Ping-Le
Piao, Yen-Ting
Chen, Hung-Wei
Hsiao, Ting-Lin
Hsu, Yun-Man
Lu, Ke-Han
Lee, Hung-yi
contents While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09714
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
Yang, Chih-Kai
Tsai, Yun-Shao
Guo, Yu-Kai
Tsai, Ping-Le
Piao, Yen-Ting
Chen, Hung-Wei
Hsiao, Ting-Lin
Hsu, Yun-Man
Lu, Ke-Han
Lee, Hung-yi
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.
title MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2603.09714