An efficient text augmentation approach for contextualized Mandarin speech recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Naijun, Wan, Xucheng, Liu, Kai, Du, Ziqing, Huan, Zhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910487485087744
author Zheng, Naijun
Wan, Xucheng
Liu, Kai
Du, Ziqing
Huan, Zhou
author_facet Zheng, Naijun
Wan, Xucheng
Liu, Kai
Du, Ziqing
Huan, Zhou
contents Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent limitations of speech-text data availability. To address this challenge, our study proposes to leverage extensive text-only datasets and contextualize pre-trained ASR models using a straightforward text-augmentation (TA) technique, all while keeping computational costs minimal. In particular, to contextualize a pre-trained CIF-based ASR, we construct a codebook using limited speech-text data. By utilizing a simple codebook lookup process, we convert available text-only data into latent text embeddings. These embeddings then enhance the inputs for the contextualized ASR. Our experiments on diverse Mandarin test sets demonstrate that our TA approach significantly boosts recognition performance. The top-performing system shows relative CER improvements of up to 30% on rare words and 15% across all words in general.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09950
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An efficient text augmentation approach for contextualized Mandarin speech recognition
Zheng, Naijun
Wan, Xucheng
Liu, Kai
Du, Ziqing
Huan, Zhou
Sound
Computation and Language
Audio and Speech Processing
Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent limitations of speech-text data availability. To address this challenge, our study proposes to leverage extensive text-only datasets and contextualize pre-trained ASR models using a straightforward text-augmentation (TA) technique, all while keeping computational costs minimal. In particular, to contextualize a pre-trained CIF-based ASR, we construct a codebook using limited speech-text data. By utilizing a simple codebook lookup process, we convert available text-only data into latent text embeddings. These embeddings then enhance the inputs for the contextualized ASR. Our experiments on diverse Mandarin test sets demonstrate that our TA approach significantly boosts recognition performance. The top-performing system shows relative CER improvements of up to 30% on rare words and 15% across all words in general.
title An efficient text augmentation approach for contextualized Mandarin speech recognition
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.09950