Noise-aware Speech Enhancement using Diffusion Probabilistic Model
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909215636848640 |
|---|---|
| author | Hu, Yuchen Chen, Chen Li, Ruizhe Zhu, Qiushi Chng, Eng Siong |
| author_facet | Hu, Yuchen Chen, Chen Li, Ruizhe Zhu, Qiushi Chng, Eng Siong |
| contents | With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2307_08029 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Noise-aware Speech Enhancement using Diffusion Probabilistic Model Hu, Yuchen Chen, Chen Li, Ruizhe Zhu, Qiushi Chng, Eng Siong Audio and Speech Processing Machine Learning Sound With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises. |
| title | Noise-aware Speech Enhancement using Diffusion Probabilistic Model |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2307.08029 |