Saved in:
Bibliographic Details
Main Authors: Plaja-Roglans, Genís, Hung, Yun-Ning, Serra, Xavier, Pereira, Igor
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.20470
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909924614733824
author Plaja-Roglans, Genís
Hung, Yun-Ning
Serra, Xavier
Pereira, Igor
author_facet Plaja-Roglans, Genís
Hung, Yun-Ning
Serra, Xavier
Pereira, Igor
contents Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
Plaja-Roglans, Genís
Hung, Yun-Ning
Serra, Xavier
Pereira, Igor
Sound
Artificial Intelligence
Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.
title Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2511.20470