DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916770446573568 |
|---|---|
| author | Chen, Xueyuan Yang, Dongchao Wu, Wenxuan Wu, Minglin Xu, Jing Wu, Xixin Wu, Zhiyong Meng, Helen |
| author_facet | Chen, Xueyuan Yang, Dongchao Wu, Wenxuan Wu, Minglin Xu, Jing Wu, Xixin Wu, Zhiyong Meng, Helen |
| contents | Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility and poor speaker similarity. In this study, we introduce a novel diffusion-based DSR system that leverages a latent diffusion model to enhance the quality of speech reconstruction. Our model comprises: (i) a speech content encoder for phoneme embedding restoration via pre-trained self-supervised learning (SSL) speech foundation models; (ii) a speaker identity encoder for speaker-aware identity preservation by in-context learning mechanism; (iii) a diffusion-based speech generator to reconstruct the speech based on the restored phoneme embedding and preserved speaker identity. Through evaluations on the widely-used UASpeech corpus, our proposed model shows notable enhancements in speech intelligibility and speaker similarity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00350 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model Chen, Xueyuan Yang, Dongchao Wu, Wenxuan Wu, Minglin Xu, Jing Wu, Xixin Wu, Zhiyong Meng, Helen Sound Audio and Speech Processing Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility and poor speaker similarity. In this study, we introduce a novel diffusion-based DSR system that leverages a latent diffusion model to enhance the quality of speech reconstruction. Our model comprises: (i) a speech content encoder for phoneme embedding restoration via pre-trained self-supervised learning (SSL) speech foundation models; (ii) a speaker identity encoder for speaker-aware identity preservation by in-context learning mechanism; (iii) a diffusion-based speech generator to reconstruct the speech based on the restored phoneme embedding and preserved speaker identity. Through evaluations on the widely-used UASpeech corpus, our proposed model shows notable enhancements in speech intelligibility and speaker similarity. |
| title | DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.00350 |