GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhaochen, Qiao, Limeng, Wan, Guanglu, Jiang, Tingting
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915936558120960
author Liu, Zhaochen
Qiao, Limeng
Wan, Guanglu
Jiang, Tingting
author_facet Liu, Zhaochen
Qiao, Limeng
Wan, Guanglu
Jiang, Tingting
contents Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foundation models, but rely on static single-layer extractions. We identify that such an approach induces a task misalignment bias: the geometric features naturally evolve towards 3D pretraining objectives, which may contradict the heterogeneous spatial demands of MLLMs, rendering any single layer fundamentally insufficient. To resolve this, we propose GeoAlign, a novel framework that dynamically aggregates multi-layer geometric features to realign with the actual demands. GeoAlign constructs a hierarchical geometric feature bank and leverages the MLLM's original visual tokens as content-aware queries to perform layer-wise sparse routing, adaptively fetching the suitable geometric features for each patch. Extensive experiments on VSI-Bench, ScanQA, and SQA3D demonstrate that our compact 4B model effectively achieves state-of-the-art performance, even outperforming larger existing MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12630
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
Liu, Zhaochen
Qiao, Limeng
Wan, Guanglu
Jiang, Tingting
Computer Vision and Pattern Recognition
Computation and Language
Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foundation models, but rely on static single-layer extractions. We identify that such an approach induces a task misalignment bias: the geometric features naturally evolve towards 3D pretraining objectives, which may contradict the heterogeneous spatial demands of MLLMs, rendering any single layer fundamentally insufficient. To resolve this, we propose GeoAlign, a novel framework that dynamically aggregates multi-layer geometric features to realign with the actual demands. GeoAlign constructs a hierarchical geometric feature bank and leverages the MLLM's original visual tokens as content-aware queries to perform layer-wise sparse routing, adaptively fetching the suitable geometric features for each patch. Extensive experiments on VSI-Bench, ScanQA, and SQA3D demonstrate that our compact 4B model effectively achieves state-of-the-art performance, even outperforming larger existing MLLMs.
title GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.12630