Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cui, Zhongjian, Cui, Chenrui, Wang, Tianrui, He, Mengnan, Shi, Hao, Ge, Meng, Gong, Caixia, Wang, Longbiao, Dang, Jianwu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910773541863424
author Cui, Zhongjian
Cui, Chenrui
Wang, Tianrui
He, Mengnan
Shi, Hao
Ge, Meng
Gong, Caixia
Wang, Longbiao
Dang, Jianwu
author_facet Cui, Zhongjian
Cui, Chenrui
Wang, Tianrui
He, Mengnan
Shi, Hao
Ge, Meng
Gong, Caixia
Wang, Longbiao
Dang, Jianwu
contents The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
Cui, Zhongjian
Cui, Chenrui
Wang, Tianrui
He, Mengnan
Shi, Hao
Ge, Meng
Gong, Caixia
Wang, Longbiao
Dang, Jianwu
Sound
Audio and Speech Processing
The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets.
title Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.02452