Token-Weighted RNN-T for Learning from Flawed Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keren, Gil, Zhou, Wei, Kalinli, Ozlem
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917705704013824
author Keren, Gil
Zhou, Wei
Kalinli, Ozlem
author_facet Keren, Gil
Zhou, Wei
Kalinli, Ozlem
contents ASR models are commonly trained with the cross-entropy criterion to increase the probability of a target token sequence. While optimizing the probability of all tokens in the target sequence is sensible, one may want to de-emphasize tokens that reflect transcription errors. In this work, we propose a novel token-weighted RNN-T criterion that augments the RNN-T objective with token-specific weights. The new objective is used for mitigating accuracy loss from transcriptions errors in the training data, which naturally appear in two settings: pseudo-labeling and human annotation errors. Experiments results show that using our method for semi-supervised learning with pseudo-labels leads to a consistent accuracy improvement, up to 38% relative. We also analyze the accuracy degradation resulting from different levels of WER in the reference transcription, and show that token-weighted RNN-T is suitable for overcoming this degradation, recovering 64%-99% of the accuracy loss.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18108
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Token-Weighted RNN-T for Learning from Flawed Data
Keren, Gil
Zhou, Wei
Kalinli, Ozlem
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
ASR models are commonly trained with the cross-entropy criterion to increase the probability of a target token sequence. While optimizing the probability of all tokens in the target sequence is sensible, one may want to de-emphasize tokens that reflect transcription errors. In this work, we propose a novel token-weighted RNN-T criterion that augments the RNN-T objective with token-specific weights. The new objective is used for mitigating accuracy loss from transcriptions errors in the training data, which naturally appear in two settings: pseudo-labeling and human annotation errors. Experiments results show that using our method for semi-supervised learning with pseudo-labels leads to a consistent accuracy improvement, up to 38% relative. We also analyze the accuracy degradation resulting from different levels of WER in the reference transcription, and show that token-weighted RNN-T is suitable for overcoming this degradation, recovering 64%-99% of the accuracy loss.
title Token-Weighted RNN-T for Learning from Flawed Data
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.18108