Thinking with Spatial Code for Physical-World Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912946179801088 |
|---|---|
| author | Chen, Jieneng Ma, Wenxin Yuan, Ruisheng Zhang, Yunzhi Wu, Jiajun Yuille, Alan |
| author_facet | Chen, Jieneng Ma, Wenxin Yuan, Ruisheng Zhang, Yunzhi Wu, Jiajun Yuille, Alan |
| contents | We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective-aware, geometrically grounded inference. As a result, our model outperforms proprietary vision-language models on VSI-Bench, setting a new state-of-the-art. Code is available at https://github.com/Beckschen/spatialcode. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_05591 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Thinking with Spatial Code for Physical-World Video Reasoning Chen, Jieneng Ma, Wenxin Yuan, Ruisheng Zhang, Yunzhi Wu, Jiajun Yuille, Alan Computer Vision and Pattern Recognition We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective-aware, geometrically grounded inference. As a result, our model outperforms proprietary vision-language models on VSI-Bench, setting a new state-of-the-art. Code is available at https://github.com/Beckschen/spatialcode. |
| title | Thinking with Spatial Code for Physical-World Video Reasoning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.05591 |