DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xueyuan, Yang, Dongchao, Wu, Wenxuan, Wu, Minglin, Xu, Jing, Wu, Xixin, Wu, Zhiyong, Meng, Helen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916770446573568
author Chen, Xueyuan
Yang, Dongchao
Wu, Wenxuan
Wu, Minglin
Xu, Jing
Wu, Xixin
Wu, Zhiyong
Meng, Helen
author_facet Chen, Xueyuan
Yang, Dongchao
Wu, Wenxuan
Wu, Minglin
Xu, Jing
Wu, Xixin
Wu, Zhiyong
Meng, Helen
contents Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility and poor speaker similarity. In this study, we introduce a novel diffusion-based DSR system that leverages a latent diffusion model to enhance the quality of speech reconstruction. Our model comprises: (i) a speech content encoder for phoneme embedding restoration via pre-trained self-supervised learning (SSL) speech foundation models; (ii) a speaker identity encoder for speaker-aware identity preservation by in-context learning mechanism; (iii) a diffusion-based speech generator to reconstruct the speech based on the restored phoneme embedding and preserved speaker identity. Through evaluations on the widely-used UASpeech corpus, our proposed model shows notable enhancements in speech intelligibility and speaker similarity.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
Chen, Xueyuan
Yang, Dongchao
Wu, Wenxuan
Wu, Minglin
Xu, Jing
Wu, Xixin
Wu, Zhiyong
Meng, Helen
Sound
Audio and Speech Processing
Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility and poor speaker similarity. In this study, we introduce a novel diffusion-based DSR system that leverages a latent diffusion model to enhance the quality of speech reconstruction. Our model comprises: (i) a speech content encoder for phoneme embedding restoration via pre-trained self-supervised learning (SSL) speech foundation models; (ii) a speaker identity encoder for speaker-aware identity preservation by in-context learning mechanism; (iii) a diffusion-based speech generator to reconstruct the speech based on the restored phoneme embedding and preserved speaker identity. Through evaluations on the widely-used UASpeech corpus, our proposed model shows notable enhancements in speech intelligibility and speaker similarity.
title DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00350