Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Hyeonseung, Yoon, Ji Won, Kim, Sungsoo, Kim, Nam Soo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917849147113472
author Lee, Hyeonseung
Yoon, Ji Won
Kim, Sungsoo
Kim, Nam Soo
author_facet Lee, Hyeonseung
Yoon, Ji Won
Kim, Sungsoo
Kim, Nam Soo
contents Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and latency. In the conventional framework, streaming transducer models are trained to maximize the likelihood function based on non-streaming recursion rules. However, this approach leads to a mismatch between training and inference, resulting in the issue of deformed likelihood and consequently suboptimal ASR accuracy. We introduce a mathematical quantification of the gap between the actual likelihood and the deformed likelihood, namely forward variable causal compensation (FoCC). We also present its estimator, FoCCE, as a solution to estimate the exact likelihood. Through experiments on the LibriSpeech dataset, we show that FoCCE training improves the accuracy of the streaming transducers.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17537
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
Lee, Hyeonseung
Yoon, Ji Won
Kim, Sungsoo
Kim, Nam Soo
Audio and Speech Processing
Machine Learning
Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and latency. In the conventional framework, streaming transducer models are trained to maximize the likelihood function based on non-streaming recursion rules. However, this approach leads to a mismatch between training and inference, resulting in the issue of deformed likelihood and consequently suboptimal ASR accuracy. We introduce a mathematical quantification of the gap between the actual likelihood and the deformed likelihood, namely forward variable causal compensation (FoCC). We also present its estimator, FoCCE, as a solution to estimate the exact likelihood. Through experiments on the LibriSpeech dataset, we show that FoCCE training improves the accuracy of the streaming transducers.
title Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2411.17537