EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zelin, Zhang, Yupu, Adhikari, Saugat, Islam, Saiful, Xiao, Tingsong, Liu, Zibo, Chen, Shigang, Yan, Da, Jiang, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatio-Temporal Partial Sensing Forecast for Long-term Traffic
by: Liu, Zibo, et al.
Published: (2024)
by: Liu, Zibo, et al.
Published: (2024)
Spatio-temporal Multivariate Time Series Forecast with Chosen Variables
by: Liu, Zibo, et al.
Published: (2025)
by: Liu, Zibo, et al.
Published: (2025)
EvaNet: Elevation-Guided Flood Extent Mapping on Earth Imagery (Extended Version)
by: Sami, Mirza Tanzim, et al.
Published: (2024)
by: Sami, Mirza Tanzim, et al.
Published: (2024)
VNU-Bench: A Benchmarking Dataset for Multi-Source Multimodal News Video Understanding
by: Liu, Zibo, et al.
Published: (2026)
by: Liu, Zibo, et al.
Published: (2026)
XTSFormer: Cross-Temporal-Scale Transformer for Irregular-Time Event Prediction in Clinical Applications
by: Xiao, Tingsong, et al.
Published: (2024)
by: Xiao, Tingsong, et al.
Published: (2024)
Temporally Detailed Hypergraph Neural ODEs for Disease Progression Modeling
by: Xiao, Tingsong, et al.
Published: (2025)
by: Xiao, Tingsong, et al.
Published: (2025)
Accelerate Coastal Ocean Circulation Model with AI Surrogate
by: Xu, Zelin, et al.
Published: (2024)
by: Xu, Zelin, et al.
Published: (2024)
DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction
by: Zhang, Yupu, et al.
Published: (2025)
by: Zhang, Yupu, et al.
Published: (2025)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Multi-View Neural Differential Equations for Continuous-Time Stream Data in Long-Term Traffic Forecasting
by: Liu, Zibo, et al.
Published: (2024)
by: Liu, Zibo, et al.
Published: (2024)
SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
by: Braham, Nassim Ait Ali, et al.
Published: (2026)
by: Braham, Nassim Ait Ali, et al.
Published: (2026)
A Survey on Uncertainty Quantification Methods for Deep Learning
by: He, Wenchong, et al.
Published: (2023)
by: He, Wenchong, et al.
Published: (2023)
A Concise Tiling Strategy for Preserving Spatial Context in Earth Observation Imagery
by: Abrahams, Ellianna, et al.
Published: (2024)
by: Abrahams, Ellianna, et al.
Published: (2024)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
by: Shiri, Fatemeh, et al.
Published: (2024)
by: Shiri, Fatemeh, et al.
Published: (2024)
OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
by: Jiang, Yulun, et al.
Published: (2025)
by: Jiang, Yulun, et al.
Published: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
Earth Science Foundation Models: From Perception to Reasoning and Discovery
by: Zhao, Xiangyu, et al.
Published: (2026)
by: Zhao, Xiangyu, et al.
Published: (2026)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling
by: Yin, Qian, et al.
Published: (2026)
by: Yin, Qian, et al.
Published: (2026)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
by: Li, Chengzu, et al.
Published: (2025)
by: Li, Chengzu, et al.
Published: (2025)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
by: Guo, Jiajie, et al.
Published: (2025)
by: Guo, Jiajie, et al.
Published: (2025)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
by: Skelic, Lejla, et al.
Published: (2025)
by: Skelic, Lejla, et al.
Published: (2025)
Limits of Spatial Imagery Reasoning in Frontier LLM Models
by: Hayashi, Sergio Y., et al.
Published: (2026)
by: Hayashi, Sergio Y., et al.
Published: (2026)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
Similar Items
-
Spatio-Temporal Partial Sensing Forecast for Long-term Traffic
by: Liu, Zibo, et al.
Published: (2024) -
Spatio-temporal Multivariate Time Series Forecast with Chosen Variables
by: Liu, Zibo, et al.
Published: (2025) -
EvaNet: Elevation-Guided Flood Extent Mapping on Earth Imagery (Extended Version)
by: Sami, Mirza Tanzim, et al.
Published: (2024) -
VNU-Bench: A Benchmarking Dataset for Multi-Source Multimodal News Video Understanding
by: Liu, Zibo, et al.
Published: (2026) -
XTSFormer: Cross-Temporal-Scale Transformer for Irregular-Time Event Prediction in Clinical Applications
by: Xiao, Tingsong, et al.
Published: (2024)