WST: Weakly Supervised Transducer for Automatic Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914140266692608 |
|---|---|
| author | Gao, Dongji Liao, Chenda Liu, Changliang Wiesner, Matthew Garcia, Leibny Paola Povey, Daniel Khudanpur, Sanjeev Wu, Jian |
| author_facet | Gao, Dongji Liao, Chenda Liu, Changliang Wiesner, Matthew Garcia, Leibny Paola Povey, Daniel Khudanpur, Sanjeev Wu, Jian |
| contents | The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_04035 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WST: Weakly Supervised Transducer for Automatic Speech Recognition Gao, Dongji Liao, Chenda Liu, Changliang Wiesner, Matthew Garcia, Leibny Paola Povey, Daniel Khudanpur, Sanjeev Wu, Jian Computation and Language The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available. |
| title | WST: Weakly Supervised Transducer for Automatic Speech Recognition |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2511.04035 |