CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Yinghao, Li, Siyou, Yu, Juntao, Benetos, Emmanouil, Maezawa, Akira
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909664254361600
author Ma, Yinghao
Li, Siyou
Yu, Juntao
Benetos, Emmanouil
Maezawa, Akira
author_facet Ma, Yinghao
Li, Siyou
Yu, Juntao
Benetos, Emmanouil
Maezawa, Akira
contents Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or multi-choice evaluations that fail to reflect the complexity of real-world music analysis. We reinterpret a broad range of traditional MIR annotations as instruction-following formats and introduce CMI-Bench, a comprehensive music instruction following benchmark designed to evaluate audio-text LLMs on a diverse set of music information retrieval (MIR) tasks. These include genre classification, emotion regression, emotion tagging, instrument classification, pitch estimation, key detection, lyrics transcription, melody extraction, vocal technique recognition, instrument performance technique detection, music tagging, music captioning, and (down)beat tracking: reflecting core challenges in MIR research. Unlike previous benchmarks, CMI-Bench adopts standardized evaluation metrics consistent with previous state-of-the-art MIR models, ensuring direct comparability with supervised approaches. We provide an evaluation toolkit supporting all open-source audio-textual LLMs, including LTU, Qwen-audio, SALMONN, MusiLingo, etc. Experiment results reveal significant performance gaps between LLMs and supervised models, along with their culture, chronological and gender bias, highlighting the potential and limitations of current models in addressing MIR tasks. CMI-Bench establishes a unified foundation for evaluating music instruction following, driving progress in music-aware LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
Ma, Yinghao
Li, Siyou
Yu, Juntao
Benetos, Emmanouil
Maezawa, Akira
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or multi-choice evaluations that fail to reflect the complexity of real-world music analysis. We reinterpret a broad range of traditional MIR annotations as instruction-following formats and introduce CMI-Bench, a comprehensive music instruction following benchmark designed to evaluate audio-text LLMs on a diverse set of music information retrieval (MIR) tasks. These include genre classification, emotion regression, emotion tagging, instrument classification, pitch estimation, key detection, lyrics transcription, melody extraction, vocal technique recognition, instrument performance technique detection, music tagging, music captioning, and (down)beat tracking: reflecting core challenges in MIR research. Unlike previous benchmarks, CMI-Bench adopts standardized evaluation metrics consistent with previous state-of-the-art MIR models, ensuring direct comparability with supervised approaches. We provide an evaluation toolkit supporting all open-source audio-textual LLMs, including LTU, Qwen-audio, SALMONN, MusiLingo, etc. Experiment results reveal significant performance gaps between LLMs and supervised models, along with their culture, chronological and gender bias, highlighting the potential and limitations of current models in addressing MIR tasks. CMI-Bench establishes a unified foundation for evaluating music instruction following, driving progress in music-aware LLMs.
title CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2506.12285