Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rangappa, Pradeep, Carofilis, Andres, Prakash, Jeena, Kumar, Shashi, Burdisso, Sergio, Madikeri, Srikanth, Villatoro-Tello, Esau, Sharma, Bidisha, Motlicek, Petr, Hacioglu, Kadri, Venkatesan, Shankar, Vyas, Saurabh, Stolcke, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914072269684736
author Rangappa, Pradeep
Carofilis, Andres
Prakash, Jeena
Kumar, Shashi
Burdisso, Sergio
Madikeri, Srikanth
Villatoro-Tello, Esau
Sharma, Bidisha
Motlicek, Petr
Hacioglu, Kadri
Venkatesan, Shankar
Vyas, Saurabh
Stolcke, Andreas
author_facet Rangappa, Pradeep
Carofilis, Andres
Prakash, Jeena
Kumar, Shashi
Burdisso, Sergio
Madikeri, Srikanth
Villatoro-Tello, Esau
Sharma, Bidisha
Motlicek, Petr
Hacioglu, Kadri
Venkatesan, Shankar
Vyas, Saurabh
Stolcke, Andreas
contents Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
Rangappa, Pradeep
Carofilis, Andres
Prakash, Jeena
Kumar, Shashi
Burdisso, Sergio
Madikeri, Srikanth
Villatoro-Tello, Esau
Sharma, Bidisha
Motlicek, Petr
Hacioglu, Kadri
Venkatesan, Shankar
Vyas, Saurabh
Stolcke, Andreas
Computation and Language
Sound
Audio and Speech Processing
Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.
title Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.03681