Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guinot, Julien, Quinton, Elio, Fazekas, György
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915080810004480
author Guinot, Julien
Quinton, Elio
Fazekas, György
author_facet Guinot, Julien
Quinton, Elio
Fazekas, György
contents Contrastive learning has proven effective in self-supervised musical representation learning, particularly for Music Information Retrieval (MIR) tasks. However, reliance on augmentation chains for contrastive view generation and the resulting learnt invariances pose challenges when different downstream tasks require sensitivity to certain musical attributes. To address this, we propose the Leave One EquiVariant (LOEV) framework, which introduces a flexible, task-adaptive approach compared to previous work by selectively preserving information about specific augmentations, allowing the model to maintain task-relevant equivariances. We demonstrate that LOEV alleviates information loss related to learned invariances, improving performance on augmentation related tasks and retrieval without sacrificing general representation quality. Furthermore, we introduce a variant of LOEV, LOEV++, which builds a disentangled latent space by design in a self-supervised manner, and enables targeted retrieval based on augmentation related attributes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18955
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
Guinot, Julien
Quinton, Elio
Fazekas, György
Sound
Audio and Speech Processing
Contrastive learning has proven effective in self-supervised musical representation learning, particularly for Music Information Retrieval (MIR) tasks. However, reliance on augmentation chains for contrastive view generation and the resulting learnt invariances pose challenges when different downstream tasks require sensitivity to certain musical attributes. To address this, we propose the Leave One EquiVariant (LOEV) framework, which introduces a flexible, task-adaptive approach compared to previous work by selectively preserving information about specific augmentations, allowing the model to maintain task-relevant equivariances. We demonstrate that LOEV alleviates information loss related to learned invariances, improving performance on augmentation related tasks and retrieval without sacrificing general representation quality. Furthermore, we introduce a variant of LOEV, LOEV++, which builds a disentangled latent space by design in a self-supervised manner, and enables targeted retrieval based on augmentation related attributes.
title Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.18955