ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911722288185344 |
|---|---|
| author | Luo, Kaiwen Lin, Liang Zhang, Yibo Aloqaily, Moayad Tao, Jialiang Wang, Dexian Zhou, Zhenhong Zhang, Junwei Wang, Kun Sun, Li Wen, Qingsong |
| author_facet | Luo, Kaiwen Lin, Liang Zhang, Yibo Aloqaily, Moayad Tao, Jialiang Wang, Dexian Zhou, Zhenhong Zhang, Junwei Wang, Kun Sun, Li Wen, Qingsong |
| contents | Although Audio Large Language Models (ALLMs) have witnessed substantial advancements, their long audio understanding capabilities remain unexplored. A plethora of benchmarks have been proposed for general audio tasks, they predominantly focus on short-form clips, leaving without a consensus on evaluating ALLMs over extended durations. This paper proposes ChronosAudio, the first multi-task benchmark tailored for long-audio understanding in ALLMs. It encompasses six major task categories and comprises 36,000 test instances totaling over 200 hours audio, stratified into short, middle, and long-form categories to comprehensively evaluate length generalization. Extensive experiments on 16 state-of-the-art models using ChronosAudio yield three critical findings: 1.Precipitous Long-Context Collapse: ALLMs exhibit a severe inability to sustain performance, with the transition from short to long contexts triggering a staggering performance degradation of over 90% in specific tasks. 2.Structural Attention Dilution: Performance degradation stems from a fundamental failure in maintaining temporal locality; attention mechanisms suffer from significant diffusion in later sequences. 3.Restorative Ceiling of Mitigation: Current strategies only offer 50% recovery. These findings reveal significant challenges in long-audio, underscoring the urgent need for approaches to achieve robust, document-level audio reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_04876 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models Luo, Kaiwen Lin, Liang Zhang, Yibo Aloqaily, Moayad Tao, Jialiang Wang, Dexian Zhou, Zhenhong Zhang, Junwei Wang, Kun Sun, Li Wen, Qingsong Sound Although Audio Large Language Models (ALLMs) have witnessed substantial advancements, their long audio understanding capabilities remain unexplored. A plethora of benchmarks have been proposed for general audio tasks, they predominantly focus on short-form clips, leaving without a consensus on evaluating ALLMs over extended durations. This paper proposes ChronosAudio, the first multi-task benchmark tailored for long-audio understanding in ALLMs. It encompasses six major task categories and comprises 36,000 test instances totaling over 200 hours audio, stratified into short, middle, and long-form categories to comprehensively evaluate length generalization. Extensive experiments on 16 state-of-the-art models using ChronosAudio yield three critical findings: 1.Precipitous Long-Context Collapse: ALLMs exhibit a severe inability to sustain performance, with the transition from short to long contexts triggering a staggering performance degradation of over 90% in specific tasks. 2.Structural Attention Dilution: Performance degradation stems from a fundamental failure in maintaining temporal locality; attention mechanisms suffer from significant diffusion in later sequences. 3.Restorative Ceiling of Mitigation: Current strategies only offer 50% recovery. These findings reveal significant challenges in long-audio, underscoring the urgent need for approaches to achieve robust, document-level audio reasoning. |
| title | ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models |
| topic | Sound |
| url | https://arxiv.org/abs/2601.04876 |