HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Zunnan, Yu, Zhentao, Zhou, Zixiang, Zhou, Jun, Jin, Xiaoyu, Hong, Fa-Ting, Ji, Xiaozhong, Zhu, Junwei, Cai, Chengfei, Tang, Shiyu, Lin, Qin, Li, Xiu, Lu, Qinglin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916661765865472
author Xu, Zunnan
Yu, Zhentao
Zhou, Zixiang
Zhou, Jun
Jin, Xiaoyu
Hong, Fa-Ting
Ji, Xiaozhong
Zhu, Junwei
Cai, Chengfei
Tang, Shiyu
Lin, Qin
Li, Xiu
Lu, Qinglin
author_facet Xu, Zunnan
Yu, Zhentao
Zhou, Zixiang
Zhou, Jun
Jin, Xiaoyu
Hong, Fa-Ting
Ji, Xiaozhong
Zhu, Junwei
Cai, Chengfei
Tang, Shiyu
Lin, Qin
Li, Xiu
Lu, Qinglin
contents We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at https://kkakkkka.github.io/HunyuanPortrait.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
Xu, Zunnan
Yu, Zhentao
Zhou, Zixiang
Zhou, Jun
Jin, Xiaoyu
Hong, Fa-Ting
Ji, Xiaozhong
Zhu, Junwei
Cai, Chengfei
Tang, Shiyu
Lin, Qin
Li, Xiu
Lu, Qinglin
Computer Vision and Pattern Recognition
We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at https://kkakkkka.github.io/HunyuanPortrait.
title HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18860