Learning Separated Representations for Instrument-based Music Similarity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hashizume, Yuka, Li, Li, Miyashita, Atsushi, Toda, Tomoki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913945387794432
author Hashizume, Yuka
Li, Li
Miyashita, Atsushi
Toda, Tomoki
author_facet Hashizume, Yuka
Li, Li
Miyashita, Atsushi
Toda, Tomoki
contents A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using multiple networks with individual instrumental signals is effective but faces the problem that using each clean instrumental signal as a query is impractical for retrieval systems and using separated instrumental signals reduces accuracy owing to artifacts. In this paper, we present instrumental-part-based music similarity learning with a single network that takes mixed signals as input instead of individual instrumental signals. Specifically, we designed a single similarity embedding space with separated subspaces for each instrument, extracted by Conditional Similarity Networks, which are trained using the triplet loss with masks. Experimental results showed that (1) the proposed method can obtain more accurate embedding representation than using individual networks using separated signals as input in the evaluation of an instrument that had low accuracy, (2) each sub-embedding space can hold the characteristics of the corresponding instrument, and (3) the selection of similar musical pieces focusing on each instrumental sound by the proposed method can obtain human acceptance, especially when focusing on timbre.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Separated Representations for Instrument-based Music Similarity
Hashizume, Yuka
Li, Li
Miyashita, Atsushi
Toda, Tomoki
Sound
Audio and Speech Processing
A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using multiple networks with individual instrumental signals is effective but faces the problem that using each clean instrumental signal as a query is impractical for retrieval systems and using separated instrumental signals reduces accuracy owing to artifacts. In this paper, we present instrumental-part-based music similarity learning with a single network that takes mixed signals as input instead of individual instrumental signals. Specifically, we designed a single similarity embedding space with separated subspaces for each instrument, extracted by Conditional Similarity Networks, which are trained using the triplet loss with masks. Experimental results showed that (1) the proposed method can obtain more accurate embedding representation than using individual networks using separated signals as input in the evaluation of an instrument that had low accuracy, (2) each sub-embedding space can hold the characteristics of the corresponding instrument, and (3) the selection of similar musical pieces focusing on each instrumental sound by the proposed method can obtain human acceptance, especially when focusing on timbre.
title Learning Separated Representations for Instrument-based Music Similarity
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.17281