VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Qilin, Jiang, Zhengkai, Xu, Chengming, Zhang, Jiangning, Wang, Yabiao, Zhang, Xinyi, Cao, Yun, Cao, Weijian, Wang, Chengjie, Fu, Yanwei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917677514096640
author Wang, Qilin
Jiang, Zhengkai
Xu, Chengming
Zhang, Jiangning
Wang, Yabiao
Zhang, Xinyi
Cao, Yun
Cao, Weijian
Wang, Chengjie
Fu, Yanwei
author_facet Wang, Qilin
Jiang, Zhengkai
Xu, Chengming
Zhang, Jiangning
Wang, Yabiao
Zhang, Xinyi
Cao, Yun
Cao, Weijian
Wang, Chengjie
Fu, Yanwei
contents Human image animation involves generating a video from a static image by following a specified pose sequence. Current approaches typically adopt a multi-stage pipeline that separately learns appearance and motion, which often leads to appearance degradation and temporal inconsistencies. To address these issues, we propose VividPose, an innovative end-to-end pipeline based on Stable Video Diffusion (SVD) that ensures superior temporal stability. To enhance the retention of human identity, we propose an identity-aware appearance controller that integrates additional facial information without compromising other appearance details such as clothing texture and background. This approach ensures that the generated videos maintain high fidelity to the identity of human subject, preserving key facial features across various poses. To accommodate diverse human body shapes and hand movements, we introduce a geometry-aware pose controller that utilizes both dense rendering maps from SMPL-X and sparse skeleton maps. This enables accurate alignment of pose and shape in the generated videos, providing a robust framework capable of handling a wide range of body shapes and dynamic hand movements. Extensive qualitative and quantitative experiments on the UBCFashion and TikTok benchmarks demonstrate that our method achieves state-of-the-art performance. Furthermore, VividPose exhibits superior generalization capabilities on our proposed in-the-wild dataset. Codes and models will be available.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18156
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
Wang, Qilin
Jiang, Zhengkai
Xu, Chengming
Zhang, Jiangning
Wang, Yabiao
Zhang, Xinyi
Cao, Yun
Cao, Weijian
Wang, Chengjie
Fu, Yanwei
Computer Vision and Pattern Recognition
Human image animation involves generating a video from a static image by following a specified pose sequence. Current approaches typically adopt a multi-stage pipeline that separately learns appearance and motion, which often leads to appearance degradation and temporal inconsistencies. To address these issues, we propose VividPose, an innovative end-to-end pipeline based on Stable Video Diffusion (SVD) that ensures superior temporal stability. To enhance the retention of human identity, we propose an identity-aware appearance controller that integrates additional facial information without compromising other appearance details such as clothing texture and background. This approach ensures that the generated videos maintain high fidelity to the identity of human subject, preserving key facial features across various poses. To accommodate diverse human body shapes and hand movements, we introduce a geometry-aware pose controller that utilizes both dense rendering maps from SMPL-X and sparse skeleton maps. This enables accurate alignment of pose and shape in the generated videos, providing a robust framework capable of handling a wide range of body shapes and dynamic hand movements. Extensive qualitative and quantitative experiments on the UBCFashion and TikTok benchmarks demonstrate that our method achieves state-of-the-art performance. Furthermore, VividPose exhibits superior generalization capabilities on our proposed in-the-wild dataset. Codes and models will be available.
title VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18156