Saved in:
Bibliographic Details
Main Authors: Gan, Qijun, Ren, Yi, Zhang, Chen, Ye, Zhenhui, Xie, Pan, Yin, Xiang, Yuan, Zehuan, Peng, Bingyue, Zhu, Jianke
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.04847
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909719913824256
author Gan, Qijun
Ren, Yi
Zhang, Chen
Ye, Zhenhui
Xie, Pan
Yin, Xiang
Yuan, Zehuan
Peng, Bingyue
Zhu, Jianke
author_facet Gan, Qijun
Ren, Yi
Zhang, Chen
Ye, Zhenhui
Xie, Pan
Yin, Xiang
Yuan, Zehuan
Peng, Bingyue
Zhu, Jianke
contents Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also rely on fixed resolution and struggle to maintain visual consistency. To address these limitations, we propose HumanDiT, a pose-guided Diffusion Transformer (DiT)-based framework trained on a large and wild dataset containing 14,000 hours of high-quality video to produce high-fidelity videos with fine-grained body rendering. Specifically, (i) HumanDiT, built on DiT, supports numerous video resolutions and variable sequence lengths, facilitating learning for long-sequence video generation; (ii) we introduce a prefix-latent reference strategy to maintain personalized characteristics across extended sequences. Furthermore, during inference, HumanDiT leverages Keypoint-DiT to generate subsequent pose sequences, facilitating video continuation from static images or existing videos. It also utilizes a Pose Adapter to enable pose transfer with given sequences. Extensive experiments demonstrate its superior performance in generating long-form, pose-accurate videos across diverse scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
Gan, Qijun
Ren, Yi
Zhang, Chen
Ye, Zhenhui
Xie, Pan
Yin, Xiang
Yuan, Zehuan
Peng, Bingyue
Zhu, Jianke
Computer Vision and Pattern Recognition
Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also rely on fixed resolution and struggle to maintain visual consistency. To address these limitations, we propose HumanDiT, a pose-guided Diffusion Transformer (DiT)-based framework trained on a large and wild dataset containing 14,000 hours of high-quality video to produce high-fidelity videos with fine-grained body rendering. Specifically, (i) HumanDiT, built on DiT, supports numerous video resolutions and variable sequence lengths, facilitating learning for long-sequence video generation; (ii) we introduce a prefix-latent reference strategy to maintain personalized characteristics across extended sequences. Furthermore, during inference, HumanDiT leverages Keypoint-DiT to generate subsequent pose sequences, facilitating video continuation from static images or existing videos. It also utilizes a Pose Adapter to enable pose transfer with given sequences. Extensive experiments demonstrate its superior performance in generating long-form, pose-accurate videos across diverse scenarios.
title HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.04847