Towards Improved Speech Recognition through Optimized Synthetic Data Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Perrin, Yanis, Boulianne, Gilles |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
por: Serrand, Coralie, et al.
Publicado: (2025)
por: Serrand, Coralie, et al.
Publicado: (2025)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
Towards Frame-level Quality Predictions of Synthetic Speech
por: Kuhlmann, Michael, et al.
Publicado: (2025)
por: Kuhlmann, Michael, et al.
Publicado: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
por: Wagner, Dominik, et al.
Publicado: (2025)
por: Wagner, Dominik, et al.
Publicado: (2025)
Using Songs to Improve Kazakh Automatic Speech Recognition
por: Yeshpanov, Rustem
Publicado: (2026)
por: Yeshpanov, Rustem
Publicado: (2026)
Group Relative Policy Optimization for Speech Recognition
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2025)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2025)
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
por: Girish, et al.
Publicado: (2025)
por: Girish, et al.
Publicado: (2025)
Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction
por: Folkertsma, Hidde, et al.
Publicado: (2026)
por: Folkertsma, Hidde, et al.
Publicado: (2026)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
por: Lin, Zhaofeng, et al.
Publicado: (2023)
por: Lin, Zhaofeng, et al.
Publicado: (2023)
Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization
por: Wu, Yihan, et al.
Publicado: (2024)
por: Wu, Yihan, et al.
Publicado: (2024)
Improving Code-Switching Speech Recognition with TTS Data Augmentation
por: Yeo, Yue Heng, et al.
Publicado: (2026)
por: Yeo, Yue Heng, et al.
Publicado: (2026)
A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
por: Tran, Minh, et al.
Publicado: (2025)
por: Tran, Minh, et al.
Publicado: (2025)
Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection
por: Liu, Yin-Long, et al.
Publicado: (2025)
por: Liu, Yin-Long, et al.
Publicado: (2025)
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
por: Liu, Hexin, et al.
Publicado: (2025)
por: Liu, Hexin, et al.
Publicado: (2025)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
por: de Groot, Dimme, et al.
Publicado: (2026)
por: de Groot, Dimme, et al.
Publicado: (2026)
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
por: Liu, Hexin, et al.
Publicado: (2024)
por: Liu, Hexin, et al.
Publicado: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
por: Yen, Hao, et al.
Publicado: (2024)
por: Yen, Hao, et al.
Publicado: (2024)
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
por: Lin, Ju, et al.
Publicado: (2024)
por: Lin, Ju, et al.
Publicado: (2024)
Chunkwise Aligners for Streaming Speech Recognition
por: Teo, Wen Shen, et al.
Publicado: (2026)
por: Teo, Wen Shen, et al.
Publicado: (2026)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
por: Gao, Yingying, et al.
Publicado: (2026)
por: Gao, Yingying, et al.
Publicado: (2026)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
por: Aronowitz, Hagai, et al.
Publicado: (2026)
por: Aronowitz, Hagai, et al.
Publicado: (2026)
Towards Real-Time Generative Speech Restoration with Flow-Matching
por: Hsieh, Tsun-An, et al.
Publicado: (2025)
por: Hsieh, Tsun-An, et al.
Publicado: (2025)
Towards a Single ASR Model That Generalizes to Disordered Speech
por: Tobin, Jimmy, et al.
Publicado: (2024)
por: Tobin, Jimmy, et al.
Publicado: (2024)
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
por: Wang, Xin, et al.
Publicado: (2025)
por: Wang, Xin, et al.
Publicado: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
por: Leung, Wing-Zin, et al.
Publicado: (2024)
por: Leung, Wing-Zin, et al.
Publicado: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
por: Cui, Mingyu, et al.
Publicado: (2025)
por: Cui, Mingyu, et al.
Publicado: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
Identifying and Calibrating Overconfidence in Noisy Speech Recognition
por: Huo, Mingyue, et al.
Publicado: (2025)
por: Huo, Mingyue, et al.
Publicado: (2025)
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
por: Su, Bo-Hao, et al.
Publicado: (2025)
por: Su, Bo-Hao, et al.
Publicado: (2025)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
In-Materia Speech Recognition
por: Zolfagharinejad, Mohamadreza, et al.
Publicado: (2024)
por: Zolfagharinejad, Mohamadreza, et al.
Publicado: (2024)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
por: Li, Zhipeng, et al.
Publicado: (2024)
por: Li, Zhipeng, et al.
Publicado: (2024)
Unsupervised Online Continual Learning for Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2024)
por: Eeckt, Steven Vander, et al.
Publicado: (2024)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
por: Dutta, Soumya, et al.
Publicado: (2024)
por: Dutta, Soumya, et al.
Publicado: (2024)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
por: Ueda, Lucas H., et al.
Publicado: (2026)
por: Ueda, Lucas H., et al.
Publicado: (2026)
Neural Encoding Detection is Not All You Need for Synthetic Speech Detection
por: Cuccovillo, Luca, et al.
Publicado: (2026)
por: Cuccovillo, Luca, et al.
Publicado: (2026)
Ejemplares similares
-
Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
por: Serrand, Coralie, et al.
Publicado: (2025) -
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
por: Bondaruk, Łukasz, et al.
Publicado: (2024) -
Towards Frame-level Quality Predictions of Synthetic Speech
por: Kuhlmann, Michael, et al.
Publicado: (2025) -
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
por: Wagner, Dominik, et al.
Publicado: (2025) -
Using Songs to Improve Kazakh Automatic Speech Recognition
por: Yeshpanov, Rustem
Publicado: (2026)