Unified Semi-Supervised Pipeline for Automatic Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910996038156288 |
|---|---|
| author | Tadevosyan, Nune Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Jukic, Ante |
| author_facet | Tadevosyan, Nune Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Jukic, Ante |
| contents | Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_07659 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unified Semi-Supervised Pipeline for Automatic Speech Recognition Tadevosyan, Nune Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Jukic, Ante Audio and Speech Processing I.5.1 Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish. |
| title | Unified Semi-Supervised Pipeline for Automatic Speech Recognition |
| topic | Audio and Speech Processing I.5.1 |
| url | https://arxiv.org/abs/2506.07659 |