CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tian-Hao, Zhou, Dinghao, Zhong, Guiping, Zhou, Jiaming, Li, Baoxiang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912134068174848
author Zhang, Tian-Hao
Zhou, Dinghao
Zhong, Guiping
Zhou, Jiaming
Li, Baoxiang
author_facet Zhang, Tian-Hao
Zhou, Dinghao
Zhong, Guiping
Zhou, Jiaming
Li, Baoxiang
contents RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor network, respectively. In this paper, we propose a novel model named CIF-Transducer (CIF-T) which incorporates the Continuous Integrate-and-Fire (CIF) mechanism with the RNN-T model to achieve efficient alignment. In this way, the RNN-T loss is abandoned, thus bringing a computational reduction and allowing the predictor network a more significant role. We also introduce Funnel-CIF, Context Blocks, Unified Gating and Bilinear Pooling joint network, and auxiliary training strategy to further improve performance. Experiments on the 178-hour AISHELL-1 and 10000-hour WenetSpeech datasets show that CIF-T achieves state-of-the-art results with lower computational overhead compared to RNN-T models.
format Preprint
id arxiv_https___arxiv_org_abs_2307_14132
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
Zhang, Tian-Hao
Zhou, Dinghao
Zhong, Guiping
Zhou, Jiaming
Li, Baoxiang
Sound
Computation and Language
Audio and Speech Processing
RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor network, respectively. In this paper, we propose a novel model named CIF-Transducer (CIF-T) which incorporates the Continuous Integrate-and-Fire (CIF) mechanism with the RNN-T model to achieve efficient alignment. In this way, the RNN-T loss is abandoned, thus bringing a computational reduction and allowing the predictor network a more significant role. We also introduce Funnel-CIF, Context Blocks, Unified Gating and Bilinear Pooling joint network, and auxiliary training strategy to further improve performance. Experiments on the 178-hour AISHELL-1 and 10000-hour WenetSpeech datasets show that CIF-T achieves state-of-the-art results with lower computational overhead compared to RNN-T models.
title CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2307.14132