Universal Music Representations? Evaluating Foundation Models on World Music Corpora

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Papaioannou, Charilaos, Benetos, Emmanouil, Potamianos, Alexandros
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915353197543424
author Papaioannou, Charilaos
Benetos, Emmanouil
Potamianos, Alexandros
author_facet Papaioannou, Charilaos
Benetos, Emmanouil
Potamianos, Alexandros
contents Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio foundation models across six musical corpora spanning Western popular, Greek, Turkish, and Indian classical traditions. We employ three complementary methodologies to investigate these models' cross-cultural capabilities: probing to assess inherent representations, targeted supervised fine-tuning of 1-2 layers, and multi-label few-shot learning for low-resource scenarios. Our analysis shows varying cross-cultural generalization, with larger models typically outperforming on non-Western music, though results decline for culturally distant traditions. Notably, our approaches achieve state-of-the-art performance on five out of six evaluated datasets, demonstrating the effectiveness of foundation models for world music understanding. We also find that our targeted fine-tuning approach does not consistently outperform probing across all settings, suggesting foundation models already encode substantial musical knowledge. Our evaluation framework and benchmarking results contribute to understanding how far current models are from achieving universal music representations while establishing metrics for future progress.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Music Representations? Evaluating Foundation Models on World Music Corpora
Papaioannou, Charilaos
Benetos, Emmanouil
Potamianos, Alexandros
Sound
Information Retrieval
Machine Learning
Audio and Speech Processing
Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio foundation models across six musical corpora spanning Western popular, Greek, Turkish, and Indian classical traditions. We employ three complementary methodologies to investigate these models' cross-cultural capabilities: probing to assess inherent representations, targeted supervised fine-tuning of 1-2 layers, and multi-label few-shot learning for low-resource scenarios. Our analysis shows varying cross-cultural generalization, with larger models typically outperforming on non-Western music, though results decline for culturally distant traditions. Notably, our approaches achieve state-of-the-art performance on five out of six evaluated datasets, demonstrating the effectiveness of foundation models for world music understanding. We also find that our targeted fine-tuning approach does not consistently outperform probing across all settings, suggesting foundation models already encode substantial musical knowledge. Our evaluation framework and benchmarking results contribute to understanding how far current models are from achieving universal music representations while establishing metrics for future progress.
title Universal Music Representations? Evaluating Foundation Models on World Music Corpora
topic Sound
Information Retrieval
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.17055