VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Sixiao, Yin, Minghao, Hu, Wenbo, Li, Xiaoyu, Shan, Ying, Fu, Yanwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914431172083712
author Zheng, Sixiao
Yin, Minghao
Hu, Wenbo
Li, Xiaoyu
Shan, Ying
Fu, Yanwei
author_facet Zheng, Sixiao
Yin, Minghao
Hu, Wenbo
Li, Xiaoyu
Shan, Ying
Fu, Yanwei
contents Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geometry-driven video world model that generates dynamic, realistic videos from a unified 4D geometric world state. Our approach is centered on a novel 4D Geometric Control representation, which encodes the world state as a static background point cloud and per-object 3D Gaussian trajectories. This representation captures each object's motion path and probabilistic 3D occupancy over time, providing a flexible, category-agnostic alternative to rigid bounding boxes and parametric models. We render 4D Geometric Control into 4D control maps for a pretrained video diffusion model, enabling high-fidelity, view-consistent video generation that faithfully follows the specified dynamics. To enable training at scale, we develop an automatic data engine and construct VerseControl4D, a real-world dataset of 35K training samples with automatically derived prompts and rendered 4D control maps. Extensive experiments show that VerseCrafter achieves superior visual quality and more accurate control over camera and multi-object motion than prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05138
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
Zheng, Sixiao
Yin, Minghao
Hu, Wenbo
Li, Xiaoyu
Shan, Ying
Fu, Yanwei
Computer Vision and Pattern Recognition
Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geometry-driven video world model that generates dynamic, realistic videos from a unified 4D geometric world state. Our approach is centered on a novel 4D Geometric Control representation, which encodes the world state as a static background point cloud and per-object 3D Gaussian trajectories. This representation captures each object's motion path and probabilistic 3D occupancy over time, providing a flexible, category-agnostic alternative to rigid bounding boxes and parametric models. We render 4D Geometric Control into 4D control maps for a pretrained video diffusion model, enabling high-fidelity, view-consistent video generation that faithfully follows the specified dynamics. To enable training at scale, we develop an automatic data engine and construct VerseControl4D, a real-world dataset of 35K training samples with automatically derived prompts and rendered 4D control maps. Extensive experiments show that VerseCrafter achieves superior visual quality and more accurate control over camera and multi-object motion than prior methods.
title VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.05138