Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Runwu, Li, Kai, Li, Chang, Wang, Jiang, Tan, Sihan, Nakadai, Kazuhiro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183550808064
author Shi, Runwu
Li, Kai
Li, Chang
Wang, Jiang
Tan, Sihan
Nakadai, Kazuhiro
author_facet Shi, Runwu
Li, Kai
Li, Chang
Wang, Jiang
Tan, Sihan
Nakadai, Kazuhiro
contents Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect real-world conditions. Instead, we revisit the source-model paradigm, training a diffusion generative model solely on anechoic speech and formulating separation as a diffusion inverse problem. However, unconditional diffusion models lack speaker-level conditioning, they can capture local acoustic structure but produce temporally inconsistent speaker identities in separated sources. To address this limitation, we propose Speaker-Embedding guidance that, during the reverse diffusion process, maintains speaker coherence within each separated track while driving embeddings of different speakers further apart. In addition, we propose a new separation-oriented solver tailored for speech separation, and both strategies effectively enhance performance on the challenging task of unsupervised source-model-based speech separation, as confirmed by extensive experimental results. Audio samples and code are available at https://runwushi.github.io/UnSepDiff_demo.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
Shi, Runwu
Li, Kai
Li, Chang
Wang, Jiang
Tan, Sihan
Nakadai, Kazuhiro
Audio and Speech Processing
Sound
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect real-world conditions. Instead, we revisit the source-model paradigm, training a diffusion generative model solely on anechoic speech and formulating separation as a diffusion inverse problem. However, unconditional diffusion models lack speaker-level conditioning, they can capture local acoustic structure but produce temporally inconsistent speaker identities in separated sources. To address this limitation, we propose Speaker-Embedding guidance that, during the reverse diffusion process, maintains speaker coherence within each separated track while driving embeddings of different speakers further apart. In addition, we propose a new separation-oriented solver tailored for speech separation, and both strategies effectively enhance performance on the challenging task of unsupervised source-model-based speech separation, as confirmed by extensive experimental results. Audio samples and code are available at https://runwushi.github.io/UnSepDiff_demo.
title Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.24395