Wan-Animate: Unified Character Animation and Replacement with Holistic Replication

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Gang, Gao, Xin, Hu, Li, Hu, Siqi, Huang, Mingyang, Ji, Chaonan, Li, Ju, Meng, Dechao, Qi, Jinwei, Qiao, Penchong, Shen, Zhen, Song, Yafei, Sun, Ke, Tian, Linrui, Wang, Feng, Wang, Guangyuan, Wang, Qi, Wang, Zhongjian, Xiao, Jiayu, Xu, Sheng, Zhang, Bang, Zhang, Peng, Zhang, Xindi, Zhang, Zhe, Zhou, Jingren, Zhuo, Lian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916954833420288
author Cheng, Gang
Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Li, Ju
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Feng
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
author_facet Cheng, Gang
Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Li, Ju
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Feng
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
contents We introduce Wan-Animate, a unified framework for character animation and replacement. Given a character image and a reference video, Wan-Animate can animate the character by precisely replicating the expressions and movements of the character in the video to generate high-fidelity character videos. Alternatively, it can integrate the animated character into the reference video to replace the original character, replicating the scene's lighting and color tone to achieve seamless environmental integration. Wan-Animate is built upon the Wan model. To adapt it for character animation tasks, we employ a modified input paradigm to differentiate between reference conditions and regions for generation. This design unifies multiple tasks into a common symbolic representation. We use spatially-aligned skeleton signals to replicate body motion and implicit facial features extracted from source images to reenact expressions, enabling the generation of character videos with high controllability and expressiveness. Furthermore, to enhance environmental integration during character replacement, we develop an auxiliary Relighting LoRA. This module preserves the character's appearance consistency while applying the appropriate environmental lighting and color tone. Experimental results demonstrate that Wan-Animate achieves state-of-the-art performance. We are committed to open-sourcing the model weights and its source code.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
Cheng, Gang
Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Li, Ju
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Feng
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
Computer Vision and Pattern Recognition
We introduce Wan-Animate, a unified framework for character animation and replacement. Given a character image and a reference video, Wan-Animate can animate the character by precisely replicating the expressions and movements of the character in the video to generate high-fidelity character videos. Alternatively, it can integrate the animated character into the reference video to replace the original character, replicating the scene's lighting and color tone to achieve seamless environmental integration. Wan-Animate is built upon the Wan model. To adapt it for character animation tasks, we employ a modified input paradigm to differentiate between reference conditions and regions for generation. This design unifies multiple tasks into a common symbolic representation. We use spatially-aligned skeleton signals to replicate body motion and implicit facial features extracted from source images to reenact expressions, enabling the generation of character videos with high controllability and expressiveness. Furthermore, to enhance environmental integration during character replacement, we develop an auxiliary Relighting LoRA. This module preserves the character's appearance consistency while applying the appropriate environmental lighting and color tone. Experimental results demonstrate that Wan-Animate achieves state-of-the-art performance. We are committed to open-sourcing the model weights and its source code.
title Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.14055