StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Shangjin, Ye, Zhichao, Liu, Jialin, Xie, Weijian, Hu, Jiaqi, Peng, Zhen, Xue, Hua, Chen, Danpeng, Wang, Xiaomeng, Yang, Lei, Wang, Nan, Liu, Haomin, Zhang, Guofeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909577803464704
author Zhai, Shangjin
Ye, Zhichao
Liu, Jialin
Xie, Weijian
Hu, Jiaqi
Peng, Zhen
Xue, Hua
Chen, Danpeng
Wang, Xiaomeng
Yang, Lei
Wang, Nan
Liu, Haomin
Zhang, Guofeng
author_facet Zhai, Shangjin
Ye, Zhichao
Liu, Jialin
Xie, Weijian
Hu, Jiaqi
Peng, Zhen
Xue, Hua
Chen, Danpeng
Wang, Xiaomeng
Yang, Lei
Wang, Nan
Liu, Haomin
Zhang, Guofeng
contents Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging. To address this, we propose StarGen, a novel framework that employs a pre-trained video diffusion model in an autoregressive manner for long-range scene generation. The generation of each video clip is conditioned on the 3D warping of spatially adjacent images and the temporally overlapping image from previously generated clips, improving spatiotemporal consistency in long-range scene generation with precise pose control. The spatiotemporal condition is compatible with various input conditions, facilitating diverse tasks, including sparse view interpolation, perpetual view generation, and layout-conditioned city generation. Quantitative and qualitative evaluations demonstrate StarGen's superior scalability, fidelity, and pose accuracy compared to state-of-the-art methods. Project page: https://zju3dv.github.io/StarGen.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
Zhai, Shangjin
Ye, Zhichao
Liu, Jialin
Xie, Weijian
Hu, Jiaqi
Peng, Zhen
Xue, Hua
Chen, Danpeng
Wang, Xiaomeng
Yang, Lei
Wang, Nan
Liu, Haomin
Zhang, Guofeng
Computer Vision and Pattern Recognition
Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging. To address this, we propose StarGen, a novel framework that employs a pre-trained video diffusion model in an autoregressive manner for long-range scene generation. The generation of each video clip is conditioned on the 3D warping of spatially adjacent images and the temporally overlapping image from previously generated clips, improving spatiotemporal consistency in long-range scene generation with precise pose control. The spatiotemporal condition is compatible with various input conditions, facilitating diverse tasks, including sparse view interpolation, perpetual view generation, and layout-conditioned city generation. Quantitative and qualitative evaluations demonstrate StarGen's superior scalability, fidelity, and pose accuracy compared to state-of-the-art methods. Project page: https://zju3dv.github.io/StarGen.
title StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.05763