MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Qihao, Cao, Yunqi, Huang, Yangyu, Leong, Hui Yi, Zhang, Fan, Yap, Kim-Hui, Hu, Wei
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915736914493440
author Zhao, Qihao
Cao, Yunqi
Huang, Yangyu
Leong, Hui Yi
Zhang, Fan
Yap, Kim-Hui
Hu, Wei
author_facet Zhao, Qihao
Cao, Yunqi
Huang, Yangyu
Leong, Hui Yi
Zhang, Fan
Yap, Kim-Hui
Hu, Wei
contents Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio, which general-purpose MLLMs often fail to handle due to insufficient perceptual grounding. We introduce MuseAgent, a music-centric multimodal agent that augments language models with structured symbolic representations derived from sheet music images and performance audio. By integrating optical music recognition and automatic music transcription modules, MuseAgent enables multi-step reasoning and interaction over fine-grained musical content. To systematically evaluate music understanding capabilities, we further propose MuseBench, a benchmark covering music theory reasoning, score interpretation, and performance-level analysis across text, image, and audio modalities. Experiments show that existing MLLMs perform poorly on these tasks, while MuseAgent achieves substantial improvements, highlighting the importance of structured multimodal grounding for interactive music understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11968
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Zhao, Qihao
Cao, Yunqi
Huang, Yangyu
Leong, Hui Yi
Zhang, Fan
Yap, Kim-Hui
Hu, Wei
Multimedia
Sound
Audio and Speech Processing
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio, which general-purpose MLLMs often fail to handle due to insufficient perceptual grounding. We introduce MuseAgent, a music-centric multimodal agent that augments language models with structured symbolic representations derived from sheet music images and performance audio. By integrating optical music recognition and automatic music transcription modules, MuseAgent enables multi-step reasoning and interaction over fine-grained musical content. To systematically evaluate music understanding capabilities, we further propose MuseBench, a benchmark covering music theory reasoning, score interpretation, and performance-level analysis across text, image, and audio modalities. Experiments show that existing MLLMs perform poorly on these tasks, while MuseAgent achieves substantial improvements, highlighting the importance of structured multimodal grounding for interactive music understanding.
title MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.11968