Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shen, Fei, Wang, Cong, Gao, Junyao, Guo, Qin, Dang, Jisheng, Tang, Jinhui, Chua, Tat-Seng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912231477739520
author Shen, Fei
Wang, Cong
Gao, Junyao
Guo, Qin
Dang, Jisheng
Tang, Jinhui
Chua, Tat-Seng
author_facet Shen, Fei
Wang, Cong
Gao, Junyao
Guo, Qin
Dang, Jisheng
Tang, Jinhui
Chua, Tat-Seng
contents Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the \textbf{M}otion-priors \textbf{C}onditional \textbf{D}iffusion \textbf{M}odel (\textbf{MCDM}), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the \textbf{TalkingFace-Wild} dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Code, models, and datasets will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09533
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
Shen, Fei
Wang, Cong
Gao, Junyao
Guo, Qin
Dang, Jisheng
Tang, Jinhui
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the \textbf{M}otion-priors \textbf{C}onditional \textbf{D}iffusion \textbf{M}odel (\textbf{MCDM}), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the \textbf{TalkingFace-Wild} dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Code, models, and datasets will be publicly available.
title Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.09533