FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Yixiang, Jiang, Fan, Wang, Chiyu, Xu, Mu, Qi, Yonggang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918180158439424
author Dai, Yixiang
Jiang, Fan
Wang, Chiyu
Xu, Mu
Qi, Yonggang
author_facet Dai, Yixiang
Jiang, Fan
Wang, Chiyu
Xu, Mu
Qi, Yonggang
contents High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative priors, current video foundation models lack explicit 3D grounding capabilities, thus being limited in both spatial consistency and their utility for downstream 3D reasoning tasks. In this work, we present FantasyWorld, a geometry-enhanced framework that augments frozen video foundation models with a trainable geometric branch, enabling joint modeling of video latents and an implicit 3D field in a single forward pass. Our approach introduces cross-branch supervision, where geometry cues guide video generation and video priors regularize 3D prediction, thus yielding consistent and generalizable 3D-aware video representations. Notably, the resulting latents from the geometric branch can potentially serve as versatile representations for downstream 3D tasks such as novel view synthesis and navigation, without requiring per-scene optimization or fine-tuning. Extensive experiments show that FantasyWorld effectively bridges video imagination and 3D perception, outperforming recent geometry-consistent baselines in multi-view coherence and style consistency. Ablation studies further confirm that these gains stem from the unified backbone and cross-branch information exchange.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21657
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
Dai, Yixiang
Jiang, Fan
Wang, Chiyu
Xu, Mu
Qi, Yonggang
Computer Vision and Pattern Recognition
High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative priors, current video foundation models lack explicit 3D grounding capabilities, thus being limited in both spatial consistency and their utility for downstream 3D reasoning tasks. In this work, we present FantasyWorld, a geometry-enhanced framework that augments frozen video foundation models with a trainable geometric branch, enabling joint modeling of video latents and an implicit 3D field in a single forward pass. Our approach introduces cross-branch supervision, where geometry cues guide video generation and video priors regularize 3D prediction, thus yielding consistent and generalizable 3D-aware video representations. Notably, the resulting latents from the geometric branch can potentially serve as versatile representations for downstream 3D tasks such as novel view synthesis and navigation, without requiring per-scene optimization or fine-tuning. Extensive experiments show that FantasyWorld effectively bridges video imagination and 3D perception, outperforming recent geometry-consistent baselines in multi-view coherence and style consistency. Ablation studies further confirm that these gains stem from the unified backbone and cross-branch information exchange.
title FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.21657