Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ma, Xingpei, Cai, Jiaran, Guan, Yuansheng, Huang, Shenneng, Zhang, Qiang, Zhang, Shunsi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914092825968640
author Ma, Xingpei
Cai, Jiaran
Guan, Yuansheng
Huang, Shenneng
Zhang, Qiang
Zhang, Shunsi
author_facet Ma, Xingpei
Cai, Jiaran
Guan, Yuansheng
Huang, Shenneng
Zhang, Qiang
Zhang, Shunsi
contents Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as inaccurate lip-sync, inappropriate head posture and the lack of fine-grained control over facial expressions. In order to introduce more face-guided conditions beyond speech audio clips, a novel two-stage training framework Playmate is proposed to generate more lifelike facial expressions and talking faces. In the first stage, we introduce a decoupled implicit 3D representation along with a meticulously designed motion-decoupled module to facilitate more accurate attribute disentanglement and generate expressive talking videos directly from audio cues. Then, in the second stage, we introduce an emotion-control module to encode emotion control information into the latent space, enabling fine-grained control over emotions and thereby achieving the ability to generate talking videos with desired emotion. Extensive experiments demonstrate that Playmate not only outperforms existing state-of-the-art methods in terms of video quality, but also exhibits strong competitiveness in lip synchronization while offering improved flexibility in controlling emotion and head pose. The code will be available at https://github.com/Playmate111/Playmate.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07203
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
Ma, Xingpei
Cai, Jiaran
Guan, Yuansheng
Huang, Shenneng
Zhang, Qiang
Zhang, Shunsi
Computer Vision and Pattern Recognition
Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as inaccurate lip-sync, inappropriate head posture and the lack of fine-grained control over facial expressions. In order to introduce more face-guided conditions beyond speech audio clips, a novel two-stage training framework Playmate is proposed to generate more lifelike facial expressions and talking faces. In the first stage, we introduce a decoupled implicit 3D representation along with a meticulously designed motion-decoupled module to facilitate more accurate attribute disentanglement and generate expressive talking videos directly from audio cues. Then, in the second stage, we introduce an emotion-control module to encode emotion control information into the latent space, enabling fine-grained control over emotions and thereby achieving the ability to generate talking videos with desired emotion. Extensive experiments demonstrate that Playmate not only outperforms existing state-of-the-art methods in terms of video quality, but also exhibits strong competitiveness in lip synchronization while offering improved flexibility in controlling emotion and head pose. The code will be available at https://github.com/Playmate111/Playmate.
title Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.07203