Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916772826841088 |
|---|---|
| author | Hajal, Karl El Hermann, Enno Hovsepyan, Sevada -Doss, Mathew Magimai. |
| author_facet | Hajal, Karl El Hermann, Enno Hovsepyan, Sevada -Doss, Mathew Magimai. |
| contents | Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our approach extends the Rhythm and Voice (RnV) conversion framework by introducing a syllable-based rhythm modeling method suited for dysarthric speech. We assess its impact on ASR by training LF-MMI models and fine-tuning Whisper on converted speech. Experiments on the Torgo corpus reveal that LF-MMI achieves significant word error rate reductions, especially for more severe cases of dysarthria, while fine-tuning Whisper on converted data has minimal effect on its performance. These results highlight the potential of unsupervised rhythm and voice conversion for dysarthric ASR. Code available at: https://github.com/idiap/RnV |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_01618 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech Hajal, Karl El Hermann, Enno Hovsepyan, Sevada -Doss, Mathew Magimai. Audio and Speech Processing Artificial Intelligence Machine Learning Sound Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our approach extends the Rhythm and Voice (RnV) conversion framework by introducing a syllable-based rhythm modeling method suited for dysarthric speech. We assess its impact on ASR by training LF-MMI models and fine-tuning Whisper on converted speech. Experiments on the Torgo corpus reveal that LF-MMI achieves significant word error rate reductions, especially for more severe cases of dysarthria, while fine-tuning Whisper on converted data has minimal effect on its performance. These results highlight the potential of unsupervised rhythm and voice conversion for dysarthric ASR. Code available at: https://github.com/idiap/RnV |
| title | Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech |
| topic | Audio and Speech Processing Artificial Intelligence Machine Learning Sound |
| url | https://arxiv.org/abs/2506.01618 |