Multi-blank Transducers for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Hainan, Jia, Fei, Majumdar, Somshubra, Watanabe, Shinji, Ginsburg, Boris
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916201729359872
author Xu, Hainan
Jia, Fei
Majumdar, Somshubra
Watanabe, Shinji
Ginsburg, Boris
author_facet Xu, Hainan
Jia, Fei
Majumdar, Somshubra
Watanabe, Shinji
Ginsburg, Boris
contents This paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional blank symbols, which consume two or more input frames when emitted. We refer to the added symbols as big blanks, and the method multi-blank RNN-T. For training multi-blank RNN-Ts, we propose a novel logit under-normalization method in order to prioritize emissions of big blanks. With experiments on multiple languages and datasets, we show that multi-blank RNN-T methods could bring relative speedups of over +90%/+139% to model inference for English Librispeech and German Multilingual Librispeech datasets, respectively. The multi-blank RNN-T method also improves ASR accuracy consistently. We will release our implementation of the method in the NeMo (https://github.com/NVIDIA/NeMo) toolkit.
format Preprint
id arxiv_https___arxiv_org_abs_2211_03541
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Multi-blank Transducers for Speech Recognition
Xu, Hainan
Jia, Fei
Majumdar, Somshubra
Watanabe, Shinji
Ginsburg, Boris
Audio and Speech Processing
Machine Learning
Sound
This paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional blank symbols, which consume two or more input frames when emitted. We refer to the added symbols as big blanks, and the method multi-blank RNN-T. For training multi-blank RNN-Ts, we propose a novel logit under-normalization method in order to prioritize emissions of big blanks. With experiments on multiple languages and datasets, we show that multi-blank RNN-T methods could bring relative speedups of over +90%/+139% to model inference for English Librispeech and German Multilingual Librispeech datasets, respectively. The multi-blank RNN-T method also improves ASR accuracy consistently. We will release our implementation of the method in the NeMo (https://github.com/NVIDIA/NeMo) toolkit.
title Multi-blank Transducers for Speech Recognition
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2211.03541