WST: Weakly Supervised Transducer for Automatic Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Dongji, Liao, Chenda, Liu, Changliang, Wiesner, Matthew, Garcia, Leibny Paola, Povey, Daniel, Khudanpur, Sanjeev, Wu, Jian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914140266692608
author Gao, Dongji
Liao, Chenda
Liu, Changliang
Wiesner, Matthew
Garcia, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
Wu, Jian
author_facet Gao, Dongji
Liao, Chenda
Liu, Changliang
Wiesner, Matthew
Garcia, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
Wu, Jian
contents The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WST: Weakly Supervised Transducer for Automatic Speech Recognition
Gao, Dongji
Liao, Chenda
Liu, Changliang
Wiesner, Matthew
Garcia, Leibny Paola
Povey, Daniel
Khudanpur, Sanjeev
Wu, Jian
Computation and Language
The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available.
title WST: Weakly Supervised Transducer for Automatic Speech Recognition
topic Computation and Language
url https://arxiv.org/abs/2511.04035