VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Xiangyu, Wang, Shijie, Zhang, Fengyi, Liu, Lin, Jia, Caiyan, Song, Ziying, Huang, Zi, Luo, Yadan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917338117308416
author Sun, Xiangyu
Wang, Shijie
Zhang, Fengyi
Liu, Lin
Jia, Caiyan
Song, Ziying
Huang, Zi
Luo, Yadan
author_facet Sun, Xiangyu
Wang, Shijie
Zhang, Fengyi
Liu, Lin
Jia, Caiyan
Song, Ziying
Huang, Zi
Luo, Yadan
contents World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain geometrically inconsistent. We present VGGT-World, a geometry world model that side-steps video generation entirely and instead forecasts the temporal evolution of frozen geometry-foundation-model (GFM) features. Concretely, we repurpose the latent tokens of a frozen VGGT as the world state and train a lightweight temporal flow transformer to autoregressively predict their future trajectory. Two technical challenges arise in this high-dimensional (d=1024) feature space: (i) standard velocity-prediction flow matching collapses, and (ii) autoregressive rollout suffers from compounding exposure bias. We address the first with a clean-target (z-prediction) parameterization that yields a substantially higher signal-to-noise ratio, and the second with a two-stage latent flow-forcing curriculum that progressively conditions the model on its own partially denoised rollouts. Experiments on KITTI, Cityscapes, and TartanAir demonstrate that VGGT-World significantly outperforms the strongest baselines in depth forecasting while running 3.6-5 times faster with only 0.43B trainable parameters, establishing frozen GFM features as an effective and efficient predictive state for 3D world modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12655
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
Sun, Xiangyu
Wang, Shijie
Zhang, Fengyi
Liu, Lin
Jia, Caiyan
Song, Ziying
Huang, Zi
Luo, Yadan
Computer Vision and Pattern Recognition
World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain geometrically inconsistent. We present VGGT-World, a geometry world model that side-steps video generation entirely and instead forecasts the temporal evolution of frozen geometry-foundation-model (GFM) features. Concretely, we repurpose the latent tokens of a frozen VGGT as the world state and train a lightweight temporal flow transformer to autoregressively predict their future trajectory. Two technical challenges arise in this high-dimensional (d=1024) feature space: (i) standard velocity-prediction flow matching collapses, and (ii) autoregressive rollout suffers from compounding exposure bias. We address the first with a clean-target (z-prediction) parameterization that yields a substantially higher signal-to-noise ratio, and the second with a two-stage latent flow-forcing curriculum that progressively conditions the model on its own partially denoised rollouts. Experiments on KITTI, Cityscapes, and TartanAir demonstrate that VGGT-World significantly outperforms the strongest baselines in depth forecasting while running 3.6-5 times faster with only 0.43B trainable parameters, establishing frozen GFM features as an effective and efficient predictive state for 3D world modeling.
title VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12655