Convoifilter: A case study of doing cocktail party speech recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thai-Binh, Waibel, Alexander
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917632161087488
author Nguyen, Thai-Binh
Waibel, Alexander
author_facet Nguyen, Thai-Binh
Waibel, Alexander
contents This paper presents an end-to-end model designed to improve automatic speech recognition (ASR) for a particular speaker in a crowded, noisy environment. The model utilizes a single-channel speech enhancement module that isolates the speaker's voice from background noise (ConVoiFilter) and an ASR module. The model can decrease ASR's word error rate (WER) from 80% to 26.4% through this approach. Typically, these two components are adjusted independently due to variations in data requirements. However, speech enhancement can create anomalies that decrease ASR efficiency. By implementing a joint fine-tuning strategy, the model can reduce the WER from 26.4% in separate tuning to 14.5% in joint tuning. We openly share our pre-trained model to foster further research hf.co/nguyenvulebinh/voice-filter.
format Preprint
id arxiv_https___arxiv_org_abs_2308_11380
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Convoifilter: A case study of doing cocktail party speech recognition
Nguyen, Thai-Binh
Waibel, Alexander
Sound
Computation and Language
Audio and Speech Processing
This paper presents an end-to-end model designed to improve automatic speech recognition (ASR) for a particular speaker in a crowded, noisy environment. The model utilizes a single-channel speech enhancement module that isolates the speaker's voice from background noise (ConVoiFilter) and an ASR module. The model can decrease ASR's word error rate (WER) from 80% to 26.4% through this approach. Typically, these two components are adjusted independently due to variations in data requirements. However, speech enhancement can create anomalies that decrease ASR efficiency. By implementing a joint fine-tuning strategy, the model can reduce the WER from 26.4% in separate tuning to 14.5% in joint tuning. We openly share our pre-trained model to foster further research hf.co/nguyenvulebinh/voice-filter.
title Convoifilter: A case study of doing cocktail party speech recognition
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2308.11380