GDiffuSE: Diffusion-based speech enhancement with noise model guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yanir, Efrayim, Burshtein, David, Gannot, Sharon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908857908854784
author Yanir, Efrayim
Burshtein, David
Gannot, Sharon
author_facet Yanir, Efrayim
Burshtein, David
Gannot, Sharon
contents This paper introduces a novel speech enhancement (SE) approach based on a denoising diffusion probabilistic model (DDPM), termed Guided diffusion for speech enhancement (GDiffuSE). In contrast to conventional methods that directly map noisy speech to clean speech, our method employs a lightweight helper model to estimate the noise distribution, which is then incorporated into the diffusion denoising process via a guidance mechanism. This design improves robustness by enabling seamless adaptation to unseen noise types and by leveraging large-scale DDPMs originally trained for speech generation in the context of SE. We evaluate our approach on noisy signals obtained by adding noise samples from the BBC sound effects database to LibriSpeech utterances, showing consistent improvements over state-of-the-art baselines under mismatched noise conditions. Examples are available at our project webpage.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GDiffuSE: Diffusion-based speech enhancement with noise model guidance
Yanir, Efrayim
Burshtein, David
Gannot, Sharon
Sound
Audio and Speech Processing
This paper introduces a novel speech enhancement (SE) approach based on a denoising diffusion probabilistic model (DDPM), termed Guided diffusion for speech enhancement (GDiffuSE). In contrast to conventional methods that directly map noisy speech to clean speech, our method employs a lightweight helper model to estimate the noise distribution, which is then incorporated into the diffusion denoising process via a guidance mechanism. This design improves robustness by enabling seamless adaptation to unseen noise types and by leveraging large-scale DDPMs originally trained for speech generation in the context of SE. We evaluate our approach on noisy signals obtained by adding noise samples from the BBC sound effects database to LibriSpeech utterances, showing consistent improvements over state-of-the-art baselines under mismatched noise conditions. Examples are available at our project webpage.
title GDiffuSE: Diffusion-based speech enhancement with noise model guidance
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.04157