DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Kangwei, Liu, Junwu, Cao, Yun, Guo, Jinlin, Yi, Xiaowei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908283146600448
author Liu, Kangwei
Liu, Junwu
Cao, Yun
Guo, Jinlin
Yi, Xiaowei
author_facet Liu, Kangwei
Liu, Junwu
Cao, Yun
Guo, Jinlin
Yi, Xiaowei
contents Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
Liu, Kangwei
Liu, Junwu
Cao, Yun
Guo, Jinlin
Yi, Xiaowei
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk.
title DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.19001