Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914072269684736 |
|---|---|
| author | Rangappa, Pradeep Carofilis, Andres Prakash, Jeena Kumar, Shashi Burdisso, Sergio Madikeri, Srikanth Villatoro-Tello, Esau Sharma, Bidisha Motlicek, Petr Hacioglu, Kadri Venkatesan, Shankar Vyas, Saurabh Stolcke, Andreas |
| author_facet | Rangappa, Pradeep Carofilis, Andres Prakash, Jeena Kumar, Shashi Burdisso, Sergio Madikeri, Srikanth Villatoro-Tello, Esau Sharma, Bidisha Motlicek, Petr Hacioglu, Kadri Venkatesan, Shankar Vyas, Saurabh Stolcke, Andreas |
| contents | Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_03681 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering Rangappa, Pradeep Carofilis, Andres Prakash, Jeena Kumar, Shashi Burdisso, Sergio Madikeri, Srikanth Villatoro-Tello, Esau Sharma, Bidisha Motlicek, Petr Hacioglu, Kadri Venkatesan, Shankar Vyas, Saurabh Stolcke, Andreas Computation and Language Sound Audio and Speech Processing Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English. |
| title | Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.03681 |