SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Di, Fei, Zhengcong, Wang, Rui, Bai, Jialin, Yu, Changqian, Fan, Mingyuan, Chen, Guibin, Wen, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929717546844160
author Qiu, Di
Fei, Zhengcong
Wang, Rui
Bai, Jialin
Yu, Changqian
Fan, Mingyuan
Chen, Guibin
Wen, Xiang
author_facet Qiu, Di
Fei, Zhengcong
Wang, Rui
Bai, Jialin
Yu, Changqian
Fan, Mingyuan
Chen, Guibin
Wen, Xiang
contents We present SkyReels-A1, a simple yet effective framework built upon video diffusion Transformer to facilitate portrait image animation. Existing methodologies still encounter issues, including identity distortion, background instability, and unrealistic facial dynamics, particularly in head-only animation scenarios. Besides, extending to accommodate diverse body proportions usually leads to visual inconsistencies or unnatural articulations. To address these challenges, SkyReels-A1 capitalizes on the strong generative capabilities of video DiT, enhancing facial motion transfer precision, identity retention, and temporal coherence. The system incorporates an expression-aware conditioning module that enables seamless video synthesis driven by expression-guided landmark inputs. Integrating the facial image-text alignment module strengthens the fusion of facial attributes with motion trajectories, reinforcing identity preservation. Additionally, SkyReels-A1 incorporates a multi-stage training paradigm to incrementally refine the correlation between expressions and motion while ensuring stable identity reproduction. Extensive empirical evaluations highlight the model's ability to produce visually coherent and compositionally diverse results, making it highly applicable to domains such as virtual avatars, remote communication, and digital media generation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10841
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
Qiu, Di
Fei, Zhengcong
Wang, Rui
Bai, Jialin
Yu, Changqian
Fan, Mingyuan
Chen, Guibin
Wen, Xiang
Computer Vision and Pattern Recognition
We present SkyReels-A1, a simple yet effective framework built upon video diffusion Transformer to facilitate portrait image animation. Existing methodologies still encounter issues, including identity distortion, background instability, and unrealistic facial dynamics, particularly in head-only animation scenarios. Besides, extending to accommodate diverse body proportions usually leads to visual inconsistencies or unnatural articulations. To address these challenges, SkyReels-A1 capitalizes on the strong generative capabilities of video DiT, enhancing facial motion transfer precision, identity retention, and temporal coherence. The system incorporates an expression-aware conditioning module that enables seamless video synthesis driven by expression-guided landmark inputs. Integrating the facial image-text alignment module strengthens the fusion of facial attributes with motion trajectories, reinforcing identity preservation. Additionally, SkyReels-A1 incorporates a multi-stage training paradigm to incrementally refine the correlation between expressions and motion while ensuring stable identity reproduction. Extensive empirical evaluations highlight the model's ability to produce visually coherent and compositionally diverse results, making it highly applicable to domains such as virtual avatars, remote communication, and digital media generation.
title SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.10841