Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yizhi, Zhu, Haina, Chen, Hangting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916752318791680
author Zhou, Yizhi
Zhu, Haina
Chen, Hangting
author_facet Zhou, Yizhi
Zhu, Haina
Chen, Hangting
contents Recently, pre-trained models for music information retrieval based on self-supervised learning (SSL) are becoming popular, showing success in various downstream tasks. However, there is limited research on the specific meanings of the encoded information and their applicability. Exploring these aspects can help us better understand their capabilities and limitations, leading to more effective use in downstream tasks. In this study, we analyze the advanced music representation model MusicFM and the newly emerged SSL model MuQ. We focus on three main aspects: (i) validating the advantages of SSL models across multiple downstream tasks, (ii) exploring the specialization of layer-wise information for different tasks, and (iii) comparing performance differences when selecting specific layers. Through this analysis, we reveal insights into the structure and potential applications of SSL models in music information retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16306
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
Zhou, Yizhi
Zhu, Haina
Chen, Hangting
Sound
Artificial Intelligence
Audio and Speech Processing
Recently, pre-trained models for music information retrieval based on self-supervised learning (SSL) are becoming popular, showing success in various downstream tasks. However, there is limited research on the specific meanings of the encoded information and their applicability. Exploring these aspects can help us better understand their capabilities and limitations, leading to more effective use in downstream tasks. In this study, we analyze the advanced music representation model MusicFM and the newly emerged SSL model MuQ. We focus on three main aspects: (i) validating the advantages of SSL models across multiple downstream tasks, (ii) exploring the specialization of layer-wise information for different tasks, and (iii) comparing performance differences when selecting specific layers. Through this analysis, we reveal insights into the structure and potential applications of SSL models in music information retrieval.
title Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.16306