CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917901327400960 |
|---|---|
| author | Wu, Shangda Wang, Yashan Yuan, Ruibin Guo, Zhancheng Tan, Xu Zhang, Ge Zhou, Monan Chen, Jing Mu, Xuefeng Gao, Yuejie Dong, Yuanliang Liu, Jiafeng Li, Xiaobing Yu, Feng Sun, Maosong |
| author_facet | Wu, Shangda Wang, Yashan Yuan, Ruibin Guo, Zhancheng Tan, Xu Zhang, Ge Zhou, Monan Chen, Jing Mu, Xuefeng Gao, Yuejie Dong, Yuanliang Liu, Jiafeng Li, Xiaobing Yu, Feng Sun, Maosong |
| contents | Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_13267 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models Wu, Shangda Wang, Yashan Yuan, Ruibin Guo, Zhancheng Tan, Xu Zhang, Ge Zhou, Monan Chen, Jing Mu, Xuefeng Gao, Yuejie Dong, Yuanliang Liu, Jiafeng Li, Xiaobing Yu, Feng Sun, Maosong Sound Computation and Language Audio and Speech Processing Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval. |
| title | CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.13267 |