LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Richter, Julius, de Oliveira, Danilo, Peer, Tal, Gerkmann, Timo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911228744433664
author Richter, Julius
de Oliveira, Danilo
Peer, Tal
Gerkmann, Timo
author_facet Richter, Julius
de Oliveira, Danilo
Peer, Tal
Gerkmann, Timo
contents We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition. These findings are also supported by a formal listening experiment.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Richter, Julius
de Oliveira, Danilo
Peer, Tal
Gerkmann, Timo
Audio and Speech Processing
Sound
We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition. These findings are also supported by a formal listening experiment.
title LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.11391