The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Carone, Brandon James, Roman, Iran R., Ripollés, Pablo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908605822795776
author Carone, Brandon James
Roman, Iran R.
Ripollés, Pablo
author_facet Carone, Brandon James
Roman, Iran R.
Ripollés, Pablo
contents Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
Carone, Brandon James
Roman, Iran R.
Ripollés, Pablo
Artificial Intelligence
Sound
Audio and Speech Processing
Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
title The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
topic Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.19055