SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Jiajie, Zhu, Qingpeng, Zeng, Jin, Wu, Xiaolong, He, Changyong, Wang, Weida |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric Attention
by: Wang, Weida, et al.
Published: (2025)
by: Wang, Weida, et al.
Published: (2025)
Towards Robust Time-of-Flight Depth Denoising with Confidence-Aware Diffusion Model
by: He, Changyong, et al.
Published: (2025)
by: He, Changyong, et al.
Published: (2025)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
by: Liu, Benlin, et al.
Published: (2024)
by: Liu, Benlin, et al.
Published: (2024)
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
by: Li, Zongzhao, et al.
Published: (2025)
by: Li, Zongzhao, et al.
Published: (2025)
Learning Proposes, Geometry Disposes: A Modular Framework for Efficient Spatial Reasoning
by: Zhu, Haichao, et al.
Published: (2026)
by: Zhu, Haichao, et al.
Published: (2026)
Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation
by: Zhang, Mingda, et al.
Published: (2025)
by: Zhang, Mingda, et al.
Published: (2025)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
by: Wu, Dongyue, et al.
Published: (2024)
by: Wu, Dongyue, et al.
Published: (2024)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
by: Li, Jiawei, et al.
Published: (2026)
by: Li, Jiawei, et al.
Published: (2026)
Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation
by: Ning, Zhenhua, et al.
Published: (2025)
by: Ning, Zhenhua, et al.
Published: (2025)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
MSSFC-Net:Enhancing Building Interpretation with Multi-Scale Spatial-Spectral Feature Collaboration
by: Huo, Dehua, et al.
Published: (2025)
by: Huo, Dehua, et al.
Published: (2025)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
SIESEF-FusionNet: Spatial Inter-correlation Enhancement and Spatially-Embedded Feature Fusion Network for LiDAR Point Cloud Semantic Segmentation
by: Chen, Jiale, et al.
Published: (2024)
by: Chen, Jiale, et al.
Published: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
Boosting Reasoning in Large Multimodal Models via Activation Replay
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
by: Zheng, Naichuan, et al.
Published: (2025)
by: Zheng, Naichuan, et al.
Published: (2025)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
by: Bai, Weimin, et al.
Published: (2025)
by: Bai, Weimin, et al.
Published: (2025)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
by: Yuan, Jiangye, et al.
Published: (2026)
by: Yuan, Jiangye, et al.
Published: (2026)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
by: Cheng, Yanchun, et al.
Published: (2026)
by: Cheng, Yanchun, et al.
Published: (2026)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
by: Dongfang, Zihao, et al.
Published: (2025)
by: Dongfang, Zihao, et al.
Published: (2025)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
by: Shiri, Fatemeh, et al.
Published: (2024)
by: Shiri, Fatemeh, et al.
Published: (2024)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
by: Shi, Jian, et al.
Published: (2026)
by: Shi, Jian, et al.
Published: (2026)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
by: Jin, Zhao, et al.
Published: (2025)
by: Jin, Zhao, et al.
Published: (2025)
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
by: Mühlematter, Dominik J., et al.
Published: (2025)
by: Mühlematter, Dominik J., et al.
Published: (2025)
Similar Items
-
Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric Attention
by: Wang, Weida, et al.
Published: (2025) -
Towards Robust Time-of-Flight Depth Denoising with Confidence-Aware Diffusion Model
by: He, Changyong, et al.
Published: (2025) -
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026) -
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026) -
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)