Unified Semi-Supervised Pipeline for Automatic Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tadevosyan, Nune, Karpov, Nikolay, Andrusenko, Andrei, Lavrukhin, Vitaly, Jukic, Ante
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910996038156288
author Tadevosyan, Nune
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Jukic, Ante
author_facet Tadevosyan, Nune
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Jukic, Ante
contents Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07659
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Semi-Supervised Pipeline for Automatic Speech Recognition
Tadevosyan, Nune
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Jukic, Ante
Audio and Speech Processing
I.5.1
Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish.
title Unified Semi-Supervised Pipeline for Automatic Speech Recognition
topic Audio and Speech Processing
I.5.1
url https://arxiv.org/abs/2506.07659