PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bang, Hayeon, Choi, Eunjin, Doh, Seungheon, Nam, Juhan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914022437158912
author Bang, Hayeon
Choi, Eunjin
Doh, Seungheon
Nam, Juhan
author_facet Bang, Hayeon
Choi, Eunjin
Doh, Seungheon
Nam, Juhan
contents Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models, predominantly trained on large-scale datasets, often struggle to captures subtle semantic distinctions within homogeneous solo piano music. Furthermore, existing piano-specific representation models are typically unimodal, failing to capture the inherently multimodal nature of piano music, expressed through audio, symbolic, and textual modalities. To address these limitations, we propose PianoBind, a piano-specific multimodal joint embedding model. We systematically investigate strategies for multi-source training and modality utilization within a joint embedding framework optimized for capturing fine-grained semantic distinctions in (1) small-scale and (2) homogeneous piano datasets. Our experimental results demonstrate that PianoBind learns multimodal representations that effectively capture subtle nuances of piano music, achieving superior text-to-music retrieval performance on in-domain and out-of-domain piano datasets compared to general-purpose music joint embedding models. Moreover, our design choices offer reusable insights for multimodal representation learning with homogeneous datasets beyond piano music.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
Bang, Hayeon
Choi, Eunjin
Doh, Seungheon
Nam, Juhan
Sound
Information Retrieval
Multimedia
Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models, predominantly trained on large-scale datasets, often struggle to captures subtle semantic distinctions within homogeneous solo piano music. Furthermore, existing piano-specific representation models are typically unimodal, failing to capture the inherently multimodal nature of piano music, expressed through audio, symbolic, and textual modalities. To address these limitations, we propose PianoBind, a piano-specific multimodal joint embedding model. We systematically investigate strategies for multi-source training and modality utilization within a joint embedding framework optimized for capturing fine-grained semantic distinctions in (1) small-scale and (2) homogeneous piano datasets. Our experimental results demonstrate that PianoBind learns multimodal representations that effectively capture subtle nuances of piano music, achieving superior text-to-music retrieval performance on in-domain and out-of-domain piano datasets compared to general-purpose music joint embedding models. Moreover, our design choices offer reusable insights for multimodal representation learning with homogeneous datasets beyond piano music.
title PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
topic Sound
Information Retrieval
Multimedia
url https://arxiv.org/abs/2509.04215