LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911228744433664 |
|---|---|
| author | Richter, Julius de Oliveira, Danilo Peer, Tal Gerkmann, Timo |
| author_facet | Richter, Julius de Oliveira, Danilo Peer, Tal Gerkmann, Timo |
| contents | We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition. These findings are also supported by a formal listening experiment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_11391 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models Richter, Julius de Oliveira, Danilo Peer, Tal Gerkmann, Timo Audio and Speech Processing Sound We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition. These findings are also supported by a formal listening experiment. |
| title | LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.11391 |