Noise-aware Speech Enhancement using Diffusion Probabilistic Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Yuchen, Chen, Chen, Li, Ruizhe, Zhu, Qiushi, Chng, Eng Siong
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909215636848640
author Hu, Yuchen
Chen, Chen
Li, Ruizhe
Zhu, Qiushi
Chng, Eng Siong
author_facet Hu, Yuchen
Chen, Chen
Li, Ruizhe
Zhu, Qiushi
Chng, Eng Siong
contents With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises.
format Preprint
id arxiv_https___arxiv_org_abs_2307_08029
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Noise-aware Speech Enhancement using Diffusion Probabilistic Model
Hu, Yuchen
Chen, Chen
Li, Ruizhe
Zhu, Qiushi
Chng, Eng Siong
Audio and Speech Processing
Machine Learning
Sound
With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises.
title Noise-aware Speech Enhancement using Diffusion Probabilistic Model
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2307.08029