Interspatial Attention for Efficient 4D Human Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Ruizhi, Xu, Yinghao, Shen, Yujun, Yang, Ceyuan, Zheng, Yang, Chen, Changan, Liu, Yebin, Wetzstein, Gordon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756235223040
author Shao, Ruizhi
Xu, Yinghao
Shen, Yujun
Yang, Ceyuan
Zheng, Yang
Chen, Changan
Liu, Yebin
Wetzstein, Gordon
author_facet Shao, Ruizhi
Xu, Yinghao
Shen, Yujun
Yang, Ceyuan
Zheng, Yang
Chen, Changan
Liu, Yebin
Wetzstein, Gordon
contents Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ template-based 3D representations or emerging video generation models but suffer from poor quality or limited consistency and identity preservation when generating individual or multiple digital humans. In this paper, we introduce a new interspatial attention (ISA) mechanism as a scalable building block for modern diffusion transformer (DiT)--based video generation models. ISA is a new type of cross attention that uses relative positional encodings tailored for the generation of human videos. Leveraging a custom-developed video variation autoencoder, we train a latent ISA-based diffusion model on a large corpus of video data. Our model achieves state-of-the-art performance for 4D human video synthesis, demonstrating remarkable motion consistency and identity preservation while providing precise control of the camera and body poses. Our code and model are publicly released at https://dsaurus.github.io/isa4d/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interspatial Attention for Efficient 4D Human Video Generation
Shao, Ruizhi
Xu, Yinghao
Shen, Yujun
Yang, Ceyuan
Zheng, Yang
Chen, Changan
Liu, Yebin
Wetzstein, Gordon
Computer Vision and Pattern Recognition
Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ template-based 3D representations or emerging video generation models but suffer from poor quality or limited consistency and identity preservation when generating individual or multiple digital humans. In this paper, we introduce a new interspatial attention (ISA) mechanism as a scalable building block for modern diffusion transformer (DiT)--based video generation models. ISA is a new type of cross attention that uses relative positional encodings tailored for the generation of human videos. Leveraging a custom-developed video variation autoencoder, we train a latent ISA-based diffusion model on a large corpus of video data. Our model achieves state-of-the-art performance for 4D human video synthesis, demonstrating remarkable motion consistency and identity preservation while providing precise control of the camera and body poses. Our code and model are publicly released at https://dsaurus.github.io/isa4d/.
title Interspatial Attention for Efficient 4D Human Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.15800