Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Wagner, Dominik, Baumann, Ilja, Engert, Natalie, Lee, Seanie, Nöth, Elmar, Riedhammer, Korbinian, Bocklet, Tobias |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models for Dysfluency Detection in Stuttered Speech
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
di: Wagner, Dominik, et al.
Pubblicazione: (2023)
di: Wagner, Dominik, et al.
Pubblicazione: (2023)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
di: Engert, Natalie, et al.
Pubblicazione: (2026)
di: Engert, Natalie, et al.
Pubblicazione: (2026)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022)
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
di: Baumann, Ilja, et al.
Pubblicazione: (2025)
di: Baumann, Ilja, et al.
Pubblicazione: (2025)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
di: Simic, Christopher, et al.
Pubblicazione: (2025)
di: Simic, Christopher, et al.
Pubblicazione: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
di: Braun, Franziska, et al.
Pubblicazione: (2026)
di: Braun, Franziska, et al.
Pubblicazione: (2026)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning
di: Agarwal, Dhruuv, et al.
Pubblicazione: (2025)
di: Agarwal, Dhruuv, et al.
Pubblicazione: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
di: Zheng, Xiuwen, et al.
Pubblicazione: (2026)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2026)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
di: Xie, Xurong, et al.
Pubblicazione: (2022)
di: Xie, Xurong, et al.
Pubblicazione: (2022)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
di: HU, Shujie, et al.
Pubblicazione: (2025)
di: HU, Shujie, et al.
Pubblicazione: (2025)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Probing Whisper for Dysarthric Speech in Detection and Assessment
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
di: Freisinger, Steffen, et al.
Pubblicazione: (2026)
di: Freisinger, Steffen, et al.
Pubblicazione: (2026)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
di: Lee, Jeehyun, et al.
Pubblicazione: (2024)
di: Lee, Jeehyun, et al.
Pubblicazione: (2024)
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
di: Ravenscroft, William, et al.
Pubblicazione: (2024)
di: Ravenscroft, William, et al.
Pubblicazione: (2024)
Infusing Acoustic Pause Context into Text-Based Dementia Assessment
di: Braun, Franziska, et al.
Pubblicazione: (2024)
di: Braun, Franziska, et al.
Pubblicazione: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
di: Choi, Yerin, et al.
Pubblicazione: (2024)
di: Choi, Yerin, et al.
Pubblicazione: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
di: Mujtaba, Dena, et al.
Pubblicazione: (2025)
di: Mujtaba, Dena, et al.
Pubblicazione: (2025)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies
di: Raja, Vishnu, et al.
Pubblicazione: (2025)
di: Raja, Vishnu, et al.
Pubblicazione: (2025)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
di: Farhadipour, Aref, et al.
Pubblicazione: (2023)
di: Farhadipour, Aref, et al.
Pubblicazione: (2023)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
di: Roy, Arnab Kumar, et al.
Pubblicazione: (2025)
di: Roy, Arnab Kumar, et al.
Pubblicazione: (2025)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Large Language Models for Dysfluency Detection in Stuttered Speech
di: Wagner, Dominik, et al.
Pubblicazione: (2024) -
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
di: Wagner, Dominik, et al.
Pubblicazione: (2024) -
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
di: Wagner, Dominik, et al.
Pubblicazione: (2023) -
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
di: Engert, Natalie, et al.
Pubblicazione: (2026) -
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022)