Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916918089220096 |
|---|---|
| author | Xiao, Qing Peng, Yingshan Zhang, PeiPei |
| author_facet | Xiao, Qing Peng, Yingshan Zhang, PeiPei |
| contents | Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18732 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database Xiao, Qing Peng, Yingshan Zhang, PeiPei Sound Artificial Intelligence Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning. |
| title | Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2508.18732 |