An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
Fuente:
arXiv
Salvato in:
| Autori principali: | Ogun, Sewade, Colotte, Vincent, Vincent, Emmanuel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
di: Ogun, Sewade
Pubblicazione: (2026)
di: Ogun, Sewade
Pubblicazione: (2026)
Performant ASR Models for Medical Entities in Accented Speech
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
di: Cui, Can, et al.
Pubblicazione: (2023)
di: Cui, Can, et al.
Pubblicazione: (2023)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
Retrieval Augmented Generation based context discovery for ASR
di: Siskos, Dimitrios, et al.
Pubblicazione: (2025)
di: Siskos, Dimitrios, et al.
Pubblicazione: (2025)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM
di: Yu, Jiawei, et al.
Pubblicazione: (2024)
di: Yu, Jiawei, et al.
Pubblicazione: (2024)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
Robust ASR Error Correction with Conservative Data Filtering
di: Udagawa, Takuma, et al.
Pubblicazione: (2024)
di: Udagawa, Takuma, et al.
Pubblicazione: (2024)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
di: Tomashenko, Natalia, et al.
Pubblicazione: (2024)
di: Tomashenko, Natalia, et al.
Pubblicazione: (2024)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
di: Kamble, Anand, et al.
Pubblicazione: (2023)
di: Kamble, Anand, et al.
Pubblicazione: (2023)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
di: Zhong, Yi, et al.
Pubblicazione: (2023)
di: Zhong, Yi, et al.
Pubblicazione: (2023)
The First VoicePrivacy Attacker Challenge Evaluation Plan
di: Tomashenko, Natalia, et al.
Pubblicazione: (2024)
di: Tomashenko, Natalia, et al.
Pubblicazione: (2024)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
di: Wei, Victor Junqiu, et al.
Pubblicazione: (2024)
di: Wei, Victor Junqiu, et al.
Pubblicazione: (2024)
RWKVTTS: Yet another TTS based on RWKV-7
di: yueyu, Lin, et al.
Pubblicazione: (2025)
di: yueyu, Lin, et al.
Pubblicazione: (2025)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
di: Song, Xingchen, et al.
Pubblicazione: (2024)
di: Song, Xingchen, et al.
Pubblicazione: (2024)
Large Language Models based ASR Error Correction for Child Conversations
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
di: Bandarupalli, Srihari, et al.
Pubblicazione: (2025)
di: Bandarupalli, Srihari, et al.
Pubblicazione: (2025)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
di: Zhou, Fangru, et al.
Pubblicazione: (2025)
di: Zhou, Fangru, et al.
Pubblicazione: (2025)
Advocating Character Error Rate for Multilingual ASR Evaluation
di: K, Thennal D, et al.
Pubblicazione: (2024)
di: K, Thennal D, et al.
Pubblicazione: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
di: Singh, Jaskaran, et al.
Pubblicazione: (2025)
di: Singh, Jaskaran, et al.
Pubblicazione: (2025)
PromptASR for contextualized ASR with controllable style
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
di: Gong, Cheng, et al.
Pubblicazione: (2024)
di: Gong, Cheng, et al.
Pubblicazione: (2024)
The Multicultural Medical Assistant: Can LLMs Improve Medical ASR Errors Across Borders?
di: Adedeji, Ayo, et al.
Pubblicazione: (2025)
di: Adedeji, Ayo, et al.
Pubblicazione: (2025)
Improving ASR Contextual Biasing with Guided Attention
di: Tang, Jiyang, et al.
Pubblicazione: (2024)
di: Tang, Jiyang, et al.
Pubblicazione: (2024)
Qwen3-TTS Technical Report
di: Hu, Hangrui, et al.
Pubblicazione: (2026)
di: Hu, Hangrui, et al.
Pubblicazione: (2026)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
di: Tawara, Naohiro, et al.
Pubblicazione: (2026)
di: Tawara, Naohiro, et al.
Pubblicazione: (2026)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
di: Anand, Srija, et al.
Pubblicazione: (2024)
di: Anand, Srija, et al.
Pubblicazione: (2024)
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
di: Wang, Xi, et al.
Pubblicazione: (2026)
di: Wang, Xi, et al.
Pubblicazione: (2026)
Building English ASR model with regional language support
di: Agrawal, Purvi, et al.
Pubblicazione: (2025)
di: Agrawal, Purvi, et al.
Pubblicazione: (2025)
Towards scalable efficient on-device ASR with transfer learning
di: Pandey, Laxmi, et al.
Pubblicazione: (2024)
di: Pandey, Laxmi, et al.
Pubblicazione: (2024)
Optimizing Byte-level Representation for End-to-end ASR
di: Hsiao, Roger, et al.
Pubblicazione: (2024)
di: Hsiao, Roger, et al.
Pubblicazione: (2024)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
di: Cheng, Yao-Fei, et al.
Pubblicazione: (2024)
di: Cheng, Yao-Fei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
di: Ogun, Sewade
Pubblicazione: (2026) -
Performant ASR Models for Medical Entities in Accented Speech
di: Afonja, Tejumade, et al.
Pubblicazione: (2024) -
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
di: Cui, Can, et al.
Pubblicazione: (2023) -
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
di: Ogun, Sewade, et al.
Pubblicazione: (2024) -
Retrieval Augmented Generation based context discovery for ASR
di: Siskos, Dimitrios, et al.
Pubblicazione: (2025)