Saved in:
Bibliographic Details
Main Authors: Lin, Daniel Chenyu, Freeman, Michael, Thickstun, John
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.05550
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916051560693760
author Lin, Daniel Chenyu
Freeman, Michael
Thickstun, John
author_facet Lin, Daniel Chenyu
Freeman, Michael
Thickstun, John
contents Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05550
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Factual Music Comprehension in Large Audio Language Models
Lin, Daniel Chenyu
Freeman, Michael
Thickstun, John
Sound
Computation and Language
Machine Learning
Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.
title Assessing Factual Music Comprehension in Large Audio Language Models
topic Sound
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.05550