Exploring Generative Error Correction for Dysarthric Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | La Quatra, Moreno, Koudounas, Alkis, Salerno, Valerio Mario, Siniscalchi, Sabato Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
MVP: Multi-source Voice Pathology detection
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
ITALIC: An Italian Intent Classification Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2023)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
Speech Analysis of Language Varieties in Italy
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies
von: Raja, Vishnu, et al.
Veröffentlicht: (2025)
von: Raja, Vishnu, et al.
Veröffentlicht: (2025)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
Full-text Error Correction for Chinese Speech Recognition with Large Language Model
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
von: Moure, Pehuén, et al.
Veröffentlicht: (2026)
von: Moure, Pehuén, et al.
Veröffentlicht: (2026)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
von: Yen, Hao, et al.
Veröffentlicht: (2025)
von: Yen, Hao, et al.
Veröffentlicht: (2025)
Multi-stage Large Language Model Correction for Speech Recognition
von: Pu, Jie, et al.
Veröffentlicht: (2023)
von: Pu, Jie, et al.
Veröffentlicht: (2023)
A Few-Shot Approach to Dysarthric Speech Intelligibility Level Classification Using Transformers
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025) -
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)