Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Zilu, Du, Jun, Lee, CHin-Hui
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910289429004288
author Guo, Zilu
Du, Jun
Lee, CHin-Hui
author_facet Guo, Zilu
Du, Jun
Lee, CHin-Hui
contents In this paper, we explore a continuous modeling approach for deep-learning-based speech enhancement, focusing on the denoising process. We use a state variable to indicate the denoising process. The starting state is noisy speech and the ending state is clean speech. The noise component in the state variable decreases with the change of the state index until the noise component is 0. During training, a UNet-like neural network learns to estimate every state variable sampled from the continuous denoising process. In testing, we introduce a controlling factor as an embedding, ranging from zero to one, to the neural network, allowing us to control the level of noise reduction. This approach enables controllable speech enhancement and is adaptable to various application scenarios. Experimental results indicate that preserving a small amount of noise in the clean target benefits speech enhancement, as evidenced by improvements in both objective speech measures and automatic speech recognition performance.
format Preprint
id arxiv_https___arxiv_org_abs_2309_09270
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
Guo, Zilu
Du, Jun
Lee, CHin-Hui
Audio and Speech Processing
Artificial Intelligence
Sound
In this paper, we explore a continuous modeling approach for deep-learning-based speech enhancement, focusing on the denoising process. We use a state variable to indicate the denoising process. The starting state is noisy speech and the ending state is clean speech. The noise component in the state variable decreases with the change of the state index until the noise component is 0. During training, a UNet-like neural network learns to estimate every state variable sampled from the continuous denoising process. In testing, we introduce a controlling factor as an embedding, ranging from zero to one, to the neural network, allowing us to control the level of noise reduction. This approach enables controllable speech enhancement and is adaptable to various application scenarios. Experimental results indicate that preserving a small amount of noise in the clean target benefits speech enhancement, as evidenced by improvements in both objective speech measures and automatic speech recognition performance.
title Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2309.09270