Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Xiaofeng, Chen, Jing, Zhang, Haitong, Xing, Menglin, Wei, Jiayi, Mu, Xuefeng, Xie, Zhongqian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915312488677376
author Pan, Xiaofeng
Chen, Jing
Zhang, Haitong
Xing, Menglin
Wei, Jiayi
Mu, Xuefeng
Xie, Zhongqian
author_facet Pan, Xiaofeng
Chen, Jing
Zhang, Haitong
Xing, Menglin
Wei, Jiayi
Mu, Xuefeng
Xie, Zhongqian
contents Recent works of music representation learning mainly focus on learning acoustic music representations with unlabeled audios or further attempt to acquire multi-modal music representations with scarce annotated audio-text pairs. They either ignore the language semantics or rely on labeled audio datasets that are difficult and expensive to create. Moreover, merely modeling semantic space usually fails to achieve satisfactory performance on music recommendation tasks since the user preference space is ignored. In this paper, we propose a novel Hierarchical Two-stage Contrastive Learning (HTCL) method that models similarity from the semantic perspective to the user perspective hierarchically to learn a comprehensive music representation bridging the gap between semantic and user preference spaces. We devise a scalable audio encoder and leverage a pre-trained BERT model as the text encoder to learn audio-text semantics via large-scale contrastive pre-training. Further, we explore a simple yet effective way to exploit interaction data from our online music platform to adapt the semantic space to user preference space via contrastive fine-tuning, which differs from previous works that follow the idea of collaborative filtering. As a result, we obtain a powerful audio encoder that not only distills language semantics from the text encoder but also models similarity in user preference space with the integrity of semantic space preserved. Experimental results on both music semantic and recommendation tasks confirm the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning
Pan, Xiaofeng
Chen, Jing
Zhang, Haitong
Xing, Menglin
Wei, Jiayi
Mu, Xuefeng
Xie, Zhongqian
Sound
Information Retrieval
Audio and Speech Processing
Recent works of music representation learning mainly focus on learning acoustic music representations with unlabeled audios or further attempt to acquire multi-modal music representations with scarce annotated audio-text pairs. They either ignore the language semantics or rely on labeled audio datasets that are difficult and expensive to create. Moreover, merely modeling semantic space usually fails to achieve satisfactory performance on music recommendation tasks since the user preference space is ignored. In this paper, we propose a novel Hierarchical Two-stage Contrastive Learning (HTCL) method that models similarity from the semantic perspective to the user perspective hierarchically to learn a comprehensive music representation bridging the gap between semantic and user preference spaces. We devise a scalable audio encoder and leverage a pre-trained BERT model as the text encoder to learn audio-text semantics via large-scale contrastive pre-training. Further, we explore a simple yet effective way to exploit interaction data from our online music platform to adapt the semantic space to user preference space via contrastive fine-tuning, which differs from previous works that follow the idea of collaborative filtering. As a result, we obtain a powerful audio encoder that not only distills language semantics from the text encoder but also models similarity in user preference space with the integrity of semantic space preserved. Experimental results on both music semantic and recommendation tasks confirm the effectiveness of our method.
title Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning
topic Sound
Information Retrieval
Audio and Speech Processing
url https://arxiv.org/abs/2505.23298