Thinking with Spatial Code for Physical-World Video Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jieneng, Ma, Wenxin, Yuan, Ruisheng, Zhang, Yunzhi, Wu, Jiajun, Yuille, Alan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912946179801088
author Chen, Jieneng
Ma, Wenxin
Yuan, Ruisheng
Zhang, Yunzhi
Wu, Jiajun
Yuille, Alan
author_facet Chen, Jieneng
Ma, Wenxin
Yuan, Ruisheng
Zhang, Yunzhi
Wu, Jiajun
Yuille, Alan
contents We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective-aware, geometrically grounded inference. As a result, our model outperforms proprietary vision-language models on VSI-Bench, setting a new state-of-the-art. Code is available at https://github.com/Beckschen/spatialcode.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05591
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking with Spatial Code for Physical-World Video Reasoning
Chen, Jieneng
Ma, Wenxin
Yuan, Ruisheng
Zhang, Yunzhi
Wu, Jiajun
Yuille, Alan
Computer Vision and Pattern Recognition
We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective-aware, geometrically grounded inference. As a result, our model outperforms proprietary vision-language models on VSI-Bench, setting a new state-of-the-art. Code is available at https://github.com/Beckschen/spatialcode.
title Thinking with Spatial Code for Physical-World Video Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.05591