Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Haoyuan, Cao, Qihang, Tang, Tao, Xiang, Kun, Guo, Zihan, Han, Jianhua, Bian, JiaWang, Xu, Hang, Liang, Xiaodan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909048103763968
author Li, Haoyuan
Cao, Qihang
Tang, Tao
Xiang, Kun
Guo, Zihan
Han, Jianhua
Bian, JiaWang
Xu, Hang
Liang, Xiaodan
author_facet Li, Haoyuan
Cao, Qihang
Tang, Tao
Xiang, Kun
Guo, Zihan
Han, Jianhua
Bian, JiaWang
Xu, Hang
Liang, Xiaodan
contents Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces semantic-geometry misalignment and redundant signals. We propose GeoThinker, a framework that shifts the paradigm from passive fusion to active perception. Instead of feature mixing, GeoThinker enables the model to selectively retrieve geometric evidence conditioned on its internal reasoning demands. GeoThinker achieves this through Spatial-Grounded Fusion applied at carefully selected VLM layers, where semantic visual priors selectively query and integrate task-relevant geometry via frame-strict cross-attention, further calibrated by Importance Gating that biases per-frame attention toward task-relevant structures. Comprehensive evaluation results show that GeoThinker sets a new state-of-the-art in spatial intelligence, achieving a peak score of 72.6 on the VSI-Bench. Furthermore, GeoThinker demonstrates robust generalization and significantly improved spatial perception across complex downstream scenarios, including embodied referring and autonomous driving. Our results indicate that the ability to actively integrate spatial structures is essential for next-generation spatial intelligence. Code can be found at https://github.com/Li-Hao-yuan/GeoThinker.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06037
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
Li, Haoyuan
Cao, Qihang
Tang, Tao
Xiang, Kun
Guo, Zihan
Han, Jianhua
Bian, JiaWang
Xu, Hang
Liang, Xiaodan
Computer Vision and Pattern Recognition
Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces semantic-geometry misalignment and redundant signals. We propose GeoThinker, a framework that shifts the paradigm from passive fusion to active perception. Instead of feature mixing, GeoThinker enables the model to selectively retrieve geometric evidence conditioned on its internal reasoning demands. GeoThinker achieves this through Spatial-Grounded Fusion applied at carefully selected VLM layers, where semantic visual priors selectively query and integrate task-relevant geometry via frame-strict cross-attention, further calibrated by Importance Gating that biases per-frame attention toward task-relevant structures. Comprehensive evaluation results show that GeoThinker sets a new state-of-the-art in spatial intelligence, achieving a peak score of 72.6 on the VSI-Bench. Furthermore, GeoThinker demonstrates robust generalization and significantly improved spatial perception across complex downstream scenarios, including embodied referring and autonomous driving. Our results indicate that the ability to actively integrate spatial structures is essential for next-generation spatial intelligence. Code can be found at https://github.com/Li-Hao-yuan/GeoThinker.
title Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.06037