Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moriya, Takafumi, Ashihara, Takanori, Mimura, Masato, Sato, Hiroshi, Matsuura, Kohei, Masumura, Ryo, Asami, Taichi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912051922731008
author Moriya, Takafumi
Ashihara, Takanori
Mimura, Masato
Sato, Hiroshi
Matsuura, Kohei
Masumura, Ryo
Asami, Taichi
author_facet Moriya, Takafumi
Ashihara, Takanori
Mimura, Masato
Sato, Hiroshi
Matsuura, Kohei
Masumura, Ryo
Asami, Taichi
contents A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel internal acoustic model (IAM) training strategy to enhance HAT-based speech recognition. IAM consists of encoder and joint networks, which are fully shared and jointly trained with HAT. This joint training not only enhances the HAT training efficiency but also encourages IAM and HAT to emit blanks synchronously which skips the more expensive non-blank computation, resulting in more effective blank thresholding for faster decoding. Experiments demonstrate that the relative error reductions of the HAT with IAM compared to the vanilla HAT are statistically significant. Moreover, we introduce dual blank thresholding, which combines both HAT- and IAM-blank thresholding and a compatible decoding algorithm. This results in a 42-75% decoding speed-up with no major performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
Moriya, Takafumi
Ashihara, Takanori
Mimura, Masato
Sato, Hiroshi
Matsuura, Kohei
Masumura, Ryo
Asami, Taichi
Audio and Speech Processing
Computation and Language
Sound
A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel internal acoustic model (IAM) training strategy to enhance HAT-based speech recognition. IAM consists of encoder and joint networks, which are fully shared and jointly trained with HAT. This joint training not only enhances the HAT training efficiency but also encourages IAM and HAT to emit blanks synchronously which skips the more expensive non-blank computation, resulting in more effective blank thresholding for faster decoding. Experiments demonstrate that the relative error reductions of the HAT with IAM compared to the vanilla HAT are statistically significant. Moreover, we introduce dual blank thresholding, which combines both HAT- and IAM-blank thresholding and a compatible decoding algorithm. This results in a 42-75% decoding speed-up with no major performance degradation.
title Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2409.20313