Source Separation for A Cappella Music

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lanzendörfer, Luca A., Pinkl, Constantin, Grötschla, Florian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908569302990848
author Lanzendörfer, Luca A.
Pinkl, Constantin
Grötschla, Florian
author_facet Lanzendörfer, Luca A.
Pinkl, Constantin
Grötschla, Florian
contents In this work, we study the task of multi-singer separation in a cappella music, where the number of active singers varies across mixtures. To address this, we use a power set-based data augmentation strategy that expands limited multi-singer datasets into exponentially more training samples. To separate singers, we introduce SepACap, an adaptation of SepReformer, a state-of-the-art speaker separation model architecture. We adapt the model with periodic activations and a composite loss function that remains effective when stems are silent, enabling robust detection and separation. Experiments on the JaCappella dataset demonstrate that our approach achieves state-of-the-art performance in both full-ensemble and subset singer separation scenarios, outperforming spectrogram-based baselines while generalizing to realistic mixtures with varying numbers of singers.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26580
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Source Separation for A Cappella Music
Lanzendörfer, Luca A.
Pinkl, Constantin
Grötschla, Florian
Sound
Machine Learning
In this work, we study the task of multi-singer separation in a cappella music, where the number of active singers varies across mixtures. To address this, we use a power set-based data augmentation strategy that expands limited multi-singer datasets into exponentially more training samples. To separate singers, we introduce SepACap, an adaptation of SepReformer, a state-of-the-art speaker separation model architecture. We adapt the model with periodic activations and a composite loss function that remains effective when stems are silent, enabling robust detection and separation. Experiments on the JaCappella dataset demonstrate that our approach achieves state-of-the-art performance in both full-ensemble and subset singer separation scenarios, outperforming spectrogram-based baselines while generalizing to realistic mixtures with varying numbers of singers.
title Source Separation for A Cappella Music
topic Sound
Machine Learning
url https://arxiv.org/abs/2509.26580