MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Ziyang, Ma, Yinghao, Zhu, Yanqiao, Yang, Chen, Chao, Yi-Wen, Xu, Ruiyang, Chen, Wenxi, Chen, Yuanzhe, Chen, Zhuo, Cong, Jian, Li, Kai, Li, Keliang, Li, Siyou, Li, Xinfeng, Li, Xiquan, Lian, Zheng, Liang, Yuzhe, Liu, Minghao, Niu, Zhikang, Wang, Tianrui, Wang, Yuping, Wang, Yuxuan, Wu, Yihao, Yang, Guanrou, Yu, Jianwei, Yuan, Ruibin, Zheng, Zhisheng, Zhou, Ziya, Zhu, Haina, Xue, Wei, Benetos, Emmanouil, Yu, Kai, Chng, Eng-Siong, Chen, Xie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912382088904704
author Ma, Ziyang
Ma, Yinghao
Zhu, Yanqiao
Yang, Chen
Chao, Yi-Wen
Xu, Ruiyang
Chen, Wenxi
Chen, Yuanzhe
Chen, Zhuo
Cong, Jian
Li, Kai
Li, Keliang
Li, Siyou
Li, Xinfeng
Li, Xiquan
Lian, Zheng
Liang, Yuzhe
Liu, Minghao
Niu, Zhikang
Wang, Tianrui
Wang, Yuping
Wang, Yuxuan
Wu, Yihao
Yang, Guanrou
Yu, Jianwei
Yuan, Ruibin
Zheng, Zhisheng
Zhou, Ziya
Zhu, Haina
Xue, Wei
Benetos, Emmanouil
Yu, Kai
Chng, Eng-Siong
Chen, Xie
author_facet Ma, Ziyang
Ma, Yinghao
Zhu, Yanqiao
Yang, Chen
Chao, Yi-Wen
Xu, Ruiyang
Chen, Wenxi
Chen, Yuanzhe
Chen, Zhuo
Cong, Jian
Li, Kai
Li, Keliang
Li, Siyou
Li, Xinfeng
Li, Xiquan
Lian, Zheng
Liang, Yuzhe
Liu, Minghao
Niu, Zhikang
Wang, Tianrui
Wang, Yuping
Wang, Yuxuan
Wu, Yihao
Yang, Guanrou
Yu, Jianwei
Yuan, Ruibin
Zheng, Zhisheng
Zhou, Ziya
Zhu, Haina
Xue, Wei
Benetos, Emmanouil
Yu, Kai
Chng, Eng-Siong
Chen, Xie
contents We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Ma, Ziyang
Ma, Yinghao
Zhu, Yanqiao
Yang, Chen
Chao, Yi-Wen
Xu, Ruiyang
Chen, Wenxi
Chen, Yuanzhe
Chen, Zhuo
Cong, Jian
Li, Kai
Li, Keliang
Li, Siyou
Li, Xinfeng
Li, Xiquan
Lian, Zheng
Liang, Yuzhe
Liu, Minghao
Niu, Zhikang
Wang, Tianrui
Wang, Yuping
Wang, Yuxuan
Wu, Yihao
Yang, Guanrou
Yu, Jianwei
Yuan, Ruibin
Zheng, Zhisheng
Zhou, Ziya
Zhu, Haina
Xue, Wei
Benetos, Emmanouil
Yu, Kai
Chng, Eng-Siong
Chen, Xie
Sound
Computation and Language
Multimedia
Audio and Speech Processing
We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.
title MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
topic Sound
Computation and Language
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2505.13032