Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911183550808064 |
|---|---|
| author | Shi, Runwu Li, Kai Li, Chang Wang, Jiang Tan, Sihan Nakadai, Kazuhiro |
| author_facet | Shi, Runwu Li, Kai Li, Chang Wang, Jiang Tan, Sihan Nakadai, Kazuhiro |
| contents | Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect real-world conditions. Instead, we revisit the source-model paradigm, training a diffusion generative model solely on anechoic speech and formulating separation as a diffusion inverse problem. However, unconditional diffusion models lack speaker-level conditioning, they can capture local acoustic structure but produce temporally inconsistent speaker identities in separated sources. To address this limitation, we propose Speaker-Embedding guidance that, during the reverse diffusion process, maintains speaker coherence within each separated track while driving embeddings of different speakers further apart. In addition, we propose a new separation-oriented solver tailored for speech separation, and both strategies effectively enhance performance on the challenging task of unsupervised source-model-based speech separation, as confirmed by extensive experimental results. Audio samples and code are available at https://runwushi.github.io/UnSepDiff_demo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_24395 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance Shi, Runwu Li, Kai Li, Chang Wang, Jiang Tan, Sihan Nakadai, Kazuhiro Audio and Speech Processing Sound Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect real-world conditions. Instead, we revisit the source-model paradigm, training a diffusion generative model solely on anechoic speech and formulating separation as a diffusion inverse problem. However, unconditional diffusion models lack speaker-level conditioning, they can capture local acoustic structure but produce temporally inconsistent speaker identities in separated sources. To address this limitation, we propose Speaker-Embedding guidance that, during the reverse diffusion process, maintains speaker coherence within each separated track while driving embeddings of different speakers further apart. In addition, we propose a new separation-oriented solver tailored for speech separation, and both strategies effectively enhance performance on the challenging task of unsupervised source-model-based speech separation, as confirmed by extensive experimental results. Audio samples and code are available at https://runwushi.github.io/UnSepDiff_demo. |
| title | Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2509.24395 |