Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhong-Qiu, Pang, Ruizhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918099063668736
author Wang, Zhong-Qiu
Pang, Ruizhe
author_facet Wang, Zhong-Qiu
Pang, Ruizhe
contents In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target speech. With this observation, we propose to leverage beamformed mixture, which has a higher SNR of the target speaker than the input mixture, as a weak supervision to train deep neural networks (DNNs) to enhance the input mixture. This way, we can train enhancement models using pairs of real-recorded mixture and its beamformed mixture, and potentially realize better generalization to real mixtures, compared with only training the models on simulated mixtures, which usually mismatch real mixtures. Evaluation results on the real-recorded CHiME-4 dataset show the effectiveness of the proposed algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
Wang, Zhong-Qiu
Pang, Ruizhe
Audio and Speech Processing
In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target speech. With this observation, we propose to leverage beamformed mixture, which has a higher SNR of the target speaker than the input mixture, as a weak supervision to train deep neural networks (DNNs) to enhance the input mixture. This way, we can train enhancement models using pairs of real-recorded mixture and its beamformed mixture, and potentially realize better generalization to real mixtures, compared with only training the models on simulated mixtures, which usually mismatch real mixtures. Evaluation results on the real-recorded CHiME-4 dataset show the effectiveness of the proposed algorithm.
title Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
topic Audio and Speech Processing
url https://arxiv.org/abs/2507.15229