Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koh, Junyoung, Lee, Jaeyun, Kim, Soo Yong, Choi, Gyu Hyeong, Koh, Jung In, Phillips, Jordan, Lee, Yeonjin, Song, Min
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911584310263808
author Koh, Junyoung
Lee, Jaeyun
Kim, Soo Yong
Choi, Gyu Hyeong
Koh, Jung In
Phillips, Jordan
Lee, Yeonjin
Song, Min
author_facet Koh, Junyoung
Lee, Jaeyun
Kim, Soo Yong
Choi, Gyu Hyeong
Koh, Jung In
Phillips, Jordan
Lee, Yeonjin
Song, Min
contents Recent work on music question answering (Music-QA) has primarily focused on single-track understanding, where models answer questions about an individual audio clip using its tags, captions, or metadata. However, listeners often describe music in comparative terms, and existing benchmarks do not systematically evaluate reasoning across multiple tracks. Building on the Jamendo-QA dataset, we introduce Jamendo-MT-QA, a dataset and benchmark for multi-track comparative question answering. From Creative Commons-licensed tracks on Jamendo, we construct 36,519 comparative QA items over 12,173 track pairs, with each pair yielding three question types: yes/no, short-answer, and sentence-level questions. We describe an LLM-assisted pipeline for generating and filtering comparative questions, and benchmark representative audio-language models using both automatic metrics and LLM-as-a-Judge evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09721
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
Koh, Junyoung
Lee, Jaeyun
Kim, Soo Yong
Choi, Gyu Hyeong
Koh, Jung In
Phillips, Jordan
Lee, Yeonjin
Song, Min
Information Retrieval
Multimedia
Sound
Recent work on music question answering (Music-QA) has primarily focused on single-track understanding, where models answer questions about an individual audio clip using its tags, captions, or metadata. However, listeners often describe music in comparative terms, and existing benchmarks do not systematically evaluate reasoning across multiple tracks. Building on the Jamendo-QA dataset, we introduce Jamendo-MT-QA, a dataset and benchmark for multi-track comparative question answering. From Creative Commons-licensed tracks on Jamendo, we construct 36,519 comparative QA items over 12,173 track pairs, with each pair yielding three question types: yes/no, short-answer, and sentence-level questions. We describe an LLM-assisted pipeline for generating and filtering comparative questions, and benchmark representative audio-language models using both automatic metrics and LLM-as-a-Judge evaluation.
title Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
topic Information Retrieval
Multimedia
Sound
url https://arxiv.org/abs/2604.09721