Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xiao, Zeqi, Zhao, Yiwei, Li, Lingxiao, Lan, Yushi, Yu, Ning, Garg, Rahul, Cooper, Roshni, Taghavi, Mohammad H., Pan, Xingang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918242930393088
author Xiao, Zeqi
Zhao, Yiwei
Li, Lingxiao
Lan, Yushi
Yu, Ning
Garg, Rahul
Cooper, Roshni
Taghavi, Mohammad H.
Pan, Xingang
author_facet Xiao, Zeqi
Zhao, Yiwei
Li, Lingxiao
Lan, Yushi
Yu, Ning
Garg, Rahul
Cooper, Roshni
Taghavi, Mohammad H.
Pan, Xingang
contents We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
Xiao, Zeqi
Zhao, Yiwei
Li, Lingxiao
Lan, Yushi
Yu, Ning
Garg, Rahul
Cooper, Roshni
Taghavi, Mohammad H.
Pan, Xingang
Computer Vision and Pattern Recognition
Artificial Intelligence
We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.
title Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.03040