Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Hainan, Chen, Zhehuai, Jia, Fei, Ginsburg, Boris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909162989944832
author Xu, Hainan
Chen, Zhehuai
Jia, Fei
Ginsburg, Boris
author_facet Xu, Hainan
Chen, Zhehuai
Jia, Fei
Ginsburg, Boris
contents This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model's decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. Instead of recognition errors being evenly spread throughout an utterance, they tend to group together, with subsequent errors often following earlier ones. Our analysis shows that PET models effectively mitigate this issue by substantially reducing the likelihood of the model generating additional errors following a prior one. Our implementation will be open-sourced with the NeMo toolkit.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04295
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
Xu, Hainan
Chen, Zhehuai
Jia, Fei
Ginsburg, Boris
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model's decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. Instead of recognition errors being evenly spread throughout an utterance, they tend to group together, with subsequent errors often following earlier ones. Our analysis shows that PET models effectively mitigate this issue by substantially reducing the likelihood of the model generating additional errors following a prior one. Our implementation will be open-sourced with the NeMo toolkit.
title Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2404.04295