Saved in:
Bibliographic Details
Main Authors: Cwitkowitz, Frank, Duan, Zhiyao
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.15569
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910343104561152
author Cwitkowitz, Frank
Duan, Zhiyao
author_facet Cwitkowitz, Frank
Duan, Zhiyao
contents Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15569
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Toward Fully Self-Supervised Multi-Pitch Estimation
Cwitkowitz, Frank
Duan, Zhiyao
Audio and Speech Processing
Machine Learning
Sound
Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.
title Toward Fully Self-Supervised Multi-Pitch Estimation
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2402.15569