Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
Fuente:
arXiv
Saved in:
| Main Authors: | Minixhofer, Christoph, Klejch, Ondrej, Bell, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025)
by: Minixhofer, Christoph, et al.
Published: (2025)
TTSDS -- Text-to-Speech Distribution Score
by: Minixhofer, Christoph, et al.
Published: (2024)
by: Minixhofer, Christoph, et al.
Published: (2024)
A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic
by: Klejch, Ondřej, et al.
Published: (2025)
by: Klejch, Ondřej, et al.
Published: (2025)
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
by: Wallbridge, Sarenne, et al.
Published: (2025)
by: Wallbridge, Sarenne, et al.
Published: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
by: Polák, Peter, et al.
Published: (2023)
by: Polák, Peter, et al.
Published: (2023)
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
by: Lathouwers, Gus, et al.
Published: (2026)
by: Lathouwers, Gus, et al.
Published: (2026)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
by: Prakash, Jeena, et al.
Published: (2025)
by: Prakash, Jeena, et al.
Published: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
by: Srivastav, Vaibhav, et al.
Published: (2025)
by: Srivastav, Vaibhav, et al.
Published: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
by: Attia, Ahmed Adel, et al.
Published: (2025)
by: Attia, Ahmed Adel, et al.
Published: (2025)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
by: Zhuo, Jianheng, et al.
Published: (2025)
by: Zhuo, Jianheng, et al.
Published: (2025)
Crossmodal ASR Error Correction with Discrete Speech Units
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
by: Jacobs, Christiaan, et al.
Published: (2025)
by: Jacobs, Christiaan, et al.
Published: (2025)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
by: Huzaifah, Muhammad, et al.
Published: (2024)
by: Huzaifah, Muhammad, et al.
Published: (2024)
Phonology-Guided Speech-to-Speech Translation for African Languages
by: Ochieng, Peter, et al.
Published: (2024)
by: Ochieng, Peter, et al.
Published: (2024)
Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
SSDM: Scalable Speech Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
CTC-Assisted LLM-Based Contextual ASR
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
by: Sedghiyeh, Nima, et al.
Published: (2025)
by: Sedghiyeh, Nima, et al.
Published: (2025)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
Enhancing ASR Performance in the Medical Domain for Dravidian Languages
by: Devarakonda, Sri Charan, et al.
Published: (2026)
by: Devarakonda, Sri Charan, et al.
Published: (2026)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
by: Song, Zheshu, et al.
Published: (2024)
by: Song, Zheshu, et al.
Published: (2024)
Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
by: Cunningham, Jay L., et al.
Published: (2025)
by: Cunningham, Jay L., et al.
Published: (2025)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
by: Yang, Cheng-Yeh, et al.
Published: (2026)
by: Yang, Cheng-Yeh, et al.
Published: (2026)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
Fun-ASR Technical Report
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
by: Lameris, Harm, et al.
Published: (2025)
by: Lameris, Harm, et al.
Published: (2025)
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
by: Taherinezhad, Fatemeh, et al.
Published: (2025)
by: Taherinezhad, Fatemeh, et al.
Published: (2025)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
by: Parulekar, Amruta, et al.
Published: (2025)
by: Parulekar, Amruta, et al.
Published: (2025)
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
by: Yusuyin, Saierdaer, et al.
Published: (2025)
by: Yusuyin, Saierdaer, et al.
Published: (2025)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
Retrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity Correction
by: An, Junjie, et al.
Published: (2026)
by: An, Junjie, et al.
Published: (2026)
Efficient Adaptation of Multilingual Models for Japanese ASR
by: Bajo, Mark, et al.
Published: (2024)
by: Bajo, Mark, et al.
Published: (2024)
Streaming Speech-to-Text Translation with a SpeechLLM
by: Parcollet, Titouan, et al.
Published: (2026)
by: Parcollet, Titouan, et al.
Published: (2026)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
by: Teleki, Maria, et al.
Published: (2025)
by: Teleki, Maria, et al.
Published: (2025)
Similar Items
-
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025) -
TTSDS -- Text-to-Speech Distribution Score
by: Minixhofer, Christoph, et al.
Published: (2024) -
A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic
by: Klejch, Ondřej, et al.
Published: (2025) -
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
by: Wallbridge, Sarenne, et al.
Published: (2025) -
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)