CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shangda, Wang, Yashan, Yuan, Ruibin, Guo, Zhancheng, Tan, Xu, Zhang, Ge, Zhou, Monan, Chen, Jing, Mu, Xuefeng, Gao, Yuejie, Dong, Yuanliang, Liu, Jiafeng, Li, Xiaobing, Yu, Feng, Sun, Maosong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917901327400960
author Wu, Shangda
Wang, Yashan
Yuan, Ruibin
Guo, Zhancheng
Tan, Xu
Zhang, Ge
Zhou, Monan
Chen, Jing
Mu, Xuefeng
Gao, Yuejie
Dong, Yuanliang
Liu, Jiafeng
Li, Xiaobing
Yu, Feng
Sun, Maosong
author_facet Wu, Shangda
Wang, Yashan
Yuan, Ruibin
Guo, Zhancheng
Tan, Xu
Zhang, Ge
Zhou, Monan
Chen, Jing
Mu, Xuefeng
Gao, Yuejie
Dong, Yuanliang
Liu, Jiafeng
Li, Xiaobing
Yu, Feng
Sun, Maosong
contents Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13267
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
Wu, Shangda
Wang, Yashan
Yuan, Ruibin
Guo, Zhancheng
Tan, Xu
Zhang, Ge
Zhou, Monan
Chen, Jing
Mu, Xuefeng
Gao, Yuejie
Dong, Yuanliang
Liu, Jiafeng
Li, Xiaobing
Yu, Feng
Sun, Maosong
Sound
Computation and Language
Audio and Speech Processing
Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval.
title CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2410.13267