Zero-Shot Duet Singing Voices Separation with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Chin-Yun, Postolache, Emilian, Rodolà, Emanuele, Fazekas, György
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913555020775424
author Yu, Chin-Yun
Postolache, Emilian
Rodolà, Emanuele
Fazekas, György
author_facet Yu, Chin-Yun
Postolache, Emilian
Rodolà, Emanuele
Fazekas, György
contents In recent studies, diffusion models have shown promise as priors for solving audio inverse problems. These models allow us to sample from the posterior distribution of a target signal given an observed signal by manipulating the diffusion process. However, when separating audio sources of the same type, such as duet singing voices, the prior learned by the diffusion process may not be sufficient to maintain the consistency of the source identity in the separated audio. For example, the singer may change from one to another occasionally. Tackling this problem will be useful for separating sources in a choir, or a mixture of multiple instruments with similar timbre, without acquiring large amounts of paired data. In this paper, we examine this problem in the context of duet singing voices separation, and propose a method to enforce the coherency of singer identity by splitting the mixture into overlapping segments and performing posterior sampling in an auto-regressive manner, conditioning on the previous segment. We evaluate the proposed method on the MedleyVox dataset and show that the proposed method outperforms the naive posterior sampling baseline. Our source code and the pre-trained model are publicly available at https://github.com/iamycy/duet-svs-diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07345
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-Shot Duet Singing Voices Separation with Diffusion Models
Yu, Chin-Yun
Postolache, Emilian
Rodolà, Emanuele
Fazekas, György
Audio and Speech Processing
Sound
In recent studies, diffusion models have shown promise as priors for solving audio inverse problems. These models allow us to sample from the posterior distribution of a target signal given an observed signal by manipulating the diffusion process. However, when separating audio sources of the same type, such as duet singing voices, the prior learned by the diffusion process may not be sufficient to maintain the consistency of the source identity in the separated audio. For example, the singer may change from one to another occasionally. Tackling this problem will be useful for separating sources in a choir, or a mixture of multiple instruments with similar timbre, without acquiring large amounts of paired data. In this paper, we examine this problem in the context of duet singing voices separation, and propose a method to enforce the coherency of singer identity by splitting the mixture into overlapping segments and performing posterior sampling in an auto-regressive manner, conditioning on the previous segment. We evaluate the proposed method on the MedleyVox dataset and show that the proposed method outperforms the naive posterior sampling baseline. Our source code and the pre-trained model are publicly available at https://github.com/iamycy/duet-svs-diffusion.
title Zero-Shot Duet Singing Voices Separation with Diffusion Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2311.07345