Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fujita, Yoto, Nugraha, Aditya Arie, Di Carlo, Diego, Bando, Yoshiaki, Fontaine, Mathieu, Yoshii, Kazuyoshi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913567960203264
author Fujita, Yoto
Nugraha, Aditya Arie
Di Carlo, Diego
Bando, Yoshiaki
Fontaine, Mathieu
Yoshii, Kazuyoshi
author_facet Fujita, Yoto
Nugraha, Aditya Arie
Di Carlo, Diego
Bando, Yoshiaki
Fontaine, Mathieu
Yoshii, Kazuyoshi
contents This paper describes speech enhancement for realtime automatic speech recognition (ASR) in real environments. A standard approach to this task is to use neural beamforming that can work efficiently in an online manner. It estimates the masks of clean dry speech from a noisy echoic mixture spectrogram with a deep neural network (DNN) and then computes a enhancement filter used for beamforming. The performance of such a supervised approach, however, is drastically degraded under mismatched conditions. This calls for run-time adaptation of the DNN. Although the ground-truth speech spectrogram required for adaptation is not available at run time, blind dereverberation and separation methods such as weighted prediction error (WPE) and fast multichannel nonnegative matrix factorization (FastMNMF) can be used for generating pseudo groundtruth data from a mixture. Based on this idea, a prior work proposed a dual-process system based on a cascade of WPE and minimum variance distortionless response (MVDR) beamforming asynchronously fine-tuned by block-online FastMNMF. To integrate the dereverberation capability into neural beamforming and make it fine-tunable at run time, we propose to use weighted power minimization distortionless response (WPD) beamforming, a unified version of WPE and minimum power distortionless response (MPDR), whose joint dereverberation and denoising filter is estimated using a DNN. We evaluated the impact of run-time adaptation under various conditions with different numbers of speakers, reverberation times, and signal-to-noise ratios (SNRs).
format Preprint
id arxiv_https___arxiv_org_abs_2410_22805
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
Fujita, Yoto
Nugraha, Aditya Arie
Di Carlo, Diego
Bando, Yoshiaki
Fontaine, Mathieu
Yoshii, Kazuyoshi
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
This paper describes speech enhancement for realtime automatic speech recognition (ASR) in real environments. A standard approach to this task is to use neural beamforming that can work efficiently in an online manner. It estimates the masks of clean dry speech from a noisy echoic mixture spectrogram with a deep neural network (DNN) and then computes a enhancement filter used for beamforming. The performance of such a supervised approach, however, is drastically degraded under mismatched conditions. This calls for run-time adaptation of the DNN. Although the ground-truth speech spectrogram required for adaptation is not available at run time, blind dereverberation and separation methods such as weighted prediction error (WPE) and fast multichannel nonnegative matrix factorization (FastMNMF) can be used for generating pseudo groundtruth data from a mixture. Based on this idea, a prior work proposed a dual-process system based on a cascade of WPE and minimum variance distortionless response (MVDR) beamforming asynchronously fine-tuned by block-online FastMNMF. To integrate the dereverberation capability into neural beamforming and make it fine-tunable at run time, we propose to use weighted power minimization distortionless response (WPD) beamforming, a unified version of WPE and minimum power distortionless response (MPDR), whose joint dereverberation and denoising filter is estimated using a DNN. We evaluated the impact of run-time adaptation under various conditions with different numbers of speakers, reverberation times, and signal-to-noise ratios (SNRs).
title Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.22805