MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917329865015296 |
|---|---|
| author | Yang, Chih-Kai Tsai, Yun-Shao Guo, Yu-Kai Tsai, Ping-Le Piao, Yen-Ting Chen, Hung-Wei Hsiao, Ting-Lin Hsu, Yun-Man Lu, Ke-Han Lee, Hung-yi |
| author_facet | Yang, Chih-Kai Tsai, Yun-Shao Guo, Yu-Kai Tsai, Ping-Le Piao, Yen-Ting Chen, Hung-Wei Hsiao, Ting-Lin Hsu, Yun-Man Lu, Ke-Han Lee, Hung-yi |
| contents | While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_09714 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models Yang, Chih-Kai Tsai, Yun-Shao Guo, Yu-Kai Tsai, Ping-Le Piao, Yen-Ting Chen, Hung-Wei Hsiao, Ting-Lin Hsu, Yun-Man Lu, Ke-Han Lee, Hung-yi Sound Artificial Intelligence Computation and Language Audio and Speech Processing While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension. |
| title | MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models |
| topic | Sound Artificial Intelligence Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2603.09714 |