Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Shangda, Zhou, Ziya, Zang, Yongyi, Zheng, Yutong, Liang, Dafang, Yuan, Ruibin, Kong, Qiuqiang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911475317080064
author Wu, Shangda
Zhou, Ziya
Zang, Yongyi
Zheng, Yutong
Liang, Dafang
Yuan, Ruibin
Kong, Qiuqiang
author_facet Wu, Shangda
Zhou, Ziya
Zang, Yongyi
Zheng, Yutong
Liang, Dafang
Yuan, Ruibin
Kong, Qiuqiang
contents We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list; 2) generating cultural-background documents for each sample in the music list via LLMs; 3) extracting key attributes from those documents; and 4) constructing multiple-choice questions probing language, region associations, mood, and thematic content. We evaluate models under four conditions and report per-language accuracy. Our findings demonstrate that even state-of-the-art audio LLMs struggle to capture subtle cultural nuances without rich textual context and exhibit systematic biases in interpreting music from different cultural traditions. The dataset is publicly available on Hugging Face to foster culturally inclusive music understanding research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00533
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
Wu, Shangda
Zhou, Ziya
Zang, Yongyi
Zheng, Yutong
Liang, Dafang
Yuan, Ruibin
Kong, Qiuqiang
Sound
Audio and Speech Processing
We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list; 2) generating cultural-background documents for each sample in the music list via LLMs; 3) extracting key attributes from those documents; and 4) constructing multiple-choice questions probing language, region associations, mood, and thematic content. We evaluate models under four conditions and report per-language accuracy. Our findings demonstrate that even state-of-the-art audio LLMs struggle to capture subtle cultural nuances without rich textual context and exhibit systematic biases in interpreting music from different cultural traditions. The dataset is publicly available on Hugging Face to foster culturally inclusive music understanding research.
title Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2603.00533