CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hui, Jiang, Changhao, Wang, Hongyu, Zhang, Ming, Sun, Jiajun, Yang, Zhixiong, Cao, Yifei, Dou, Shihan, Fan, Xiaoran, Fan, Baoyu, Ji, Tao, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909982489837568
author Li, Hui
Jiang, Changhao
Wang, Hongyu
Zhang, Ming
Sun, Jiajun
Yang, Zhixiong
Cao, Yifei
Dou, Shihan
Fan, Xiaoran
Fan, Baoyu
Ji, Tao
Gui, Tao
Zhang, Qi
Huang, Xuanjing
author_facet Li, Hui
Jiang, Changhao
Wang, Hongyu
Zhang, Ming
Sun, Jiajun
Yang, Zhixiong
Cao, Yifei
Dou, Shihan
Fan, Xiaoran
Fan, Baoyu
Ji, Tao
Gui, Tao
Zhang, Qi
Huang, Xuanjing
contents The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English audio data and do not fully capture scenarios where multiple speakers, unfolding events, and heterogeneous audio sources interact. To address these challenges, we introduce CMDAR, a Chinese benchmark for evaluating models on complex, multi-scene, and dynamically evolving audio reasoning tasks. CMDAR comprises 3,000 carefully curated question-answer pairs linked to diverse audio clips, covering five categories of complex reasoning and spanning three question types. We benchmark 26 state-of-the-art audio language models on CMDAR and observe that they exhibit limitations in complex reasoning tasks. In CMDAR-main, Qwen2.5-Omni achieves 76.67% accuracy, whereas GPT-4o Audio reaches 68.47%. However, GPT-4o Audio substantially outperforms Qwen2.5-Omni on the more challenging multiple-choice with multiple audios and open-ended tasks. And we provide detail analysis corresponding suggestions for the future development of large audio language models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22461
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
Li, Hui
Jiang, Changhao
Wang, Hongyu
Zhang, Ming
Sun, Jiajun
Yang, Zhixiong
Cao, Yifei
Dou, Shihan
Fan, Xiaoran
Fan, Baoyu
Ji, Tao
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English audio data and do not fully capture scenarios where multiple speakers, unfolding events, and heterogeneous audio sources interact. To address these challenges, we introduce CMDAR, a Chinese benchmark for evaluating models on complex, multi-scene, and dynamically evolving audio reasoning tasks. CMDAR comprises 3,000 carefully curated question-answer pairs linked to diverse audio clips, covering five categories of complex reasoning and spanning three question types. We benchmark 26 state-of-the-art audio language models on CMDAR and observe that they exhibit limitations in complex reasoning tasks. In CMDAR-main, Qwen2.5-Omni achieves 76.67% accuracy, whereas GPT-4o Audio reaches 68.47%. However, GPT-4o Audio substantially outperforms Qwen2.5-Omni on the more challenging multiple-choice with multiple audios and open-ended tasks. And we provide detail analysis corresponding suggestions for the future development of large audio language models.
title CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.22461