Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Ruizhi, Pang, Youxin, Zheng, Zerong, Sun, Jingxiang, Liu, Yebin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916407988453376
author Shao, Ruizhi
Pang, Youxin
Zheng, Zerong
Sun, Jingxiang
Liu, Yebin
author_facet Shao, Ruizhi
Pang, Youxin
Zheng, Zerong
Sun, Jingxiang
Liu, Yebin
contents We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across viewpoints and time, and CNNs for accurate condition injection. The core is a hierarchical 4D transformer architecture that factorizes self-attention across views, time steps, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we collect a multi-dimensional dataset spanning images, videos, multi-view data, and limited 4D footage, along with a tailored multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on generative adversarial networks or vanilla diffusion models, which struggle with complex motions, viewpoint changes, and generalization. Through extensive experiments, we demonstrate our method's ability to synthesize 360-degree realistic, coherent human motion videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17405
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
Shao, Ruizhi
Pang, Youxin
Zheng, Zerong
Sun, Jingxiang
Liu, Yebin
Computer Vision and Pattern Recognition
We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across viewpoints and time, and CNNs for accurate condition injection. The core is a hierarchical 4D transformer architecture that factorizes self-attention across views, time steps, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we collect a multi-dimensional dataset spanning images, videos, multi-view data, and limited 4D footage, along with a tailored multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on generative adversarial networks or vanilla diffusion models, which struggle with complex motions, viewpoint changes, and generalization. Through extensive experiments, we demonstrate our method's ability to synthesize 360-degree realistic, coherent human motion videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation.
title Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.17405