Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhogale, Kaushal Santosh, Mehendale, Deovrat, Parasa, Niharika, G, Sathish Kumar Reddy, Javed, Tahir, Kumar, Pratyush, Khapra, Mitesh M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
von: Joshi, Sakshi, et al.
Veröffentlicht: (2025)
von: Joshi, Sakshi, et al.
Veröffentlicht: (2025)
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
von: Bhogale, Kaushal, et al.
Veröffentlicht: (2026)
von: Bhogale, Kaushal, et al.
Veröffentlicht: (2026)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2026)
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2026)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
von: Rangappa, Pradeep, et al.
Veröffentlicht: (2025)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
von: Anand, Srija, et al.
Veröffentlicht: (2024)
von: Anand, Srija, et al.
Veröffentlicht: (2024)
Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
Learning When to Trust Which Teacher for Weakly Supervised ASR
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2023)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2023)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
A two-stage transliteration approach to improve performance of a multilingual ASR
von: Kumar, Rohit
Veröffentlicht: (2024)
von: Kumar, Rohit
Veröffentlicht: (2024)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
von: Javed, Tahir, et al.
Veröffentlicht: (2024)
von: Javed, Tahir, et al.
Veröffentlicht: (2024)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
von: Tsutsumi, Ayuto, et al.
Veröffentlicht: (2026)
von: Tsutsumi, Ayuto, et al.
Veröffentlicht: (2026)
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Index-ASR Technical Report
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
von: Joshi, Sakshi, et al.
Veröffentlicht: (2025) -
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
von: Bhogale, Kaushal, et al.
Veröffentlicht: (2026) -
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2026) -
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025) -
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)