Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Qing, Peng, Yingshan, Zhang, PeiPei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025)
por: Singh, Satwinder, et al.
Publicado: (2025)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
por: Choi, Yerin, et al.
Publicado: (2024)
por: Choi, Yerin, et al.
Publicado: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
por: Hu, Shujie, et al.
Publicado: (2024)
por: Hu, Shujie, et al.
Publicado: (2024)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
por: Zhong, Tao, et al.
Publicado: (2025)
por: Zhong, Tao, et al.
Publicado: (2025)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
por: Aboeitta, Ahmed, et al.
Publicado: (2025)
por: Aboeitta, Ahmed, et al.
Publicado: (2025)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
por: Shayaninasab, Minoo, et al.
Publicado: (2024)
por: Shayaninasab, Minoo, et al.
Publicado: (2024)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
por: Das, Shoutrik, et al.
Publicado: (2025)
por: Das, Shoutrik, et al.
Publicado: (2025)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
CDSD: Chinese Dysarthria Speech Database
por: Wang, Yan, et al.
Publicado: (2023)
por: Wang, Yan, et al.
Publicado: (2023)
PTS-SNN: A Prompt-Tuned Temporal Shift Spiking Neural Networks for Efficient Speech Emotion Recognition
por: Su, Xun, et al.
Publicado: (2026)
por: Su, Xun, et al.
Publicado: (2026)
MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
por: Yu, Xinyue, et al.
Publicado: (2025)
por: Yu, Xinyue, et al.
Publicado: (2025)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
por: Wagner, Dominik, et al.
Publicado: (2025)
por: Wagner, Dominik, et al.
Publicado: (2025)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
por: Huang, Wenbin, et al.
Publicado: (2026)
por: Huang, Wenbin, et al.
Publicado: (2026)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
por: Bhuiyan, Mohammed Aman, et al.
Publicado: (2026)
por: Bhuiyan, Mohammed Aman, et al.
Publicado: (2026)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
por: Pillai, Leena G, et al.
Publicado: (2024)
por: Pillai, Leena G, et al.
Publicado: (2024)
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
por: Chen, Jingyi, et al.
Publicado: (2025)
por: Chen, Jingyi, et al.
Publicado: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
por: Jiao, Xinxin, et al.
Publicado: (2024)
por: Jiao, Xinxin, et al.
Publicado: (2024)
ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
por: Yu, Xincheng, et al.
Publicado: (2023)
por: Yu, Xincheng, et al.
Publicado: (2023)
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
por: Wiafe, Isaac, et al.
Publicado: (2026)
por: Wiafe, Isaac, et al.
Publicado: (2026)
Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms
por: Penumajji, Niketa
Publicado: (2025)
por: Penumajji, Niketa
Publicado: (2025)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
por: Siriwardhana, Shamane, et al.
Publicado: (2020)
por: Siriwardhana, Shamane, et al.
Publicado: (2020)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
por: Lou, Haowei, et al.
Publicado: (2024)
por: Lou, Haowei, et al.
Publicado: (2024)
Speech Emotion Recognition via Entropy-Aware Score Selection
por: Chua, ChenYi, et al.
Publicado: (2025)
por: Chua, ChenYi, et al.
Publicado: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
por: Leung, Wing-Zin, et al.
Publicado: (2024)
por: Leung, Wing-Zin, et al.
Publicado: (2024)
Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
por: Yang, Hongli, et al.
Publicado: (2025)
por: Yang, Hongli, et al.
Publicado: (2025)
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
por: Sailor, Hardik B., et al.
Publicado: (2025)
por: Sailor, Hardik B., et al.
Publicado: (2025)
HNote: Extending YNote with Hexadecimal Encoding for Fine-Tuning LLMs in Music Modeling
por: Chu, Hung-Ying, et al.
Publicado: (2025)
por: Chu, Hung-Ying, et al.
Publicado: (2025)
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
por: Božík, Erik, et al.
Publicado: (2025)
por: Božík, Erik, et al.
Publicado: (2025)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
GSRM: Generative Speech Reward Model for Speech RLHF
por: Shen, Maohao, et al.
Publicado: (2026)
por: Shen, Maohao, et al.
Publicado: (2026)
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
por: Rakotoarivony, Lucas
Publicado: (2026)
por: Rakotoarivony, Lucas
Publicado: (2026)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
por: Qiang, Chunyu, et al.
Publicado: (2024)
por: Qiang, Chunyu, et al.
Publicado: (2024)
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
por: Tiwari, Upasana, et al.
Publicado: (2025)
por: Tiwari, Upasana, et al.
Publicado: (2025)
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
por: Li, Dongyuan, et al.
Publicado: (2024)
por: Li, Dongyuan, et al.
Publicado: (2024)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
por: Kim, Jaeyoung, et al.
Publicado: (2024)
por: Kim, Jaeyoung, et al.
Publicado: (2024)
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
por: Shi, Jiacheng, et al.
Publicado: (2025)
por: Shi, Jiacheng, et al.
Publicado: (2025)
Chord Recognition with Deep Learning
por: Mackenzie, Pierre
Publicado: (2025)
por: Mackenzie, Pierre
Publicado: (2025)
Ejemplares similares
-
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025) -
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
por: Choi, Yerin, et al.
Publicado: (2024) -
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
por: Hu, Shujie, et al.
Publicado: (2024) -
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
por: Zhong, Tao, et al.
Publicado: (2025) -
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
por: Aboeitta, Ahmed, et al.
Publicado: (2025)