Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Hoehing, Nils, Maniparambil, Mayug, Rushe, Ellen, O'Connor, Noel E., Ventresque, Anthony |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Metamorphic Testing for Pose Estimation Systems
by: Duran, Matias, et al.
Published: (2025)
by: Duran, Matias, et al.
Published: (2025)
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?
by: Hashmi, Anam, et al.
Published: (2026)
by: Hashmi, Anam, et al.
Published: (2026)
Ensemble Learning with Sparse Hypercolumns
by: Dietlmeier, Julia, et al.
Published: (2026)
by: Dietlmeier, Julia, et al.
Published: (2026)
Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
by: Sirotkin, Kirill, et al.
Published: (2024)
by: Sirotkin, Kirill, et al.
Published: (2024)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
Understanding the Effect of Knowledge Graph Extraction Error on Downstream Graph Analyses: A Case Study on Affiliation Graphs
by: Cai, Erica, et al.
Published: (2025)
by: Cai, Erica, et al.
Published: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
by: Dan, Nifu, et al.
Published: (2025)
by: Dan, Nifu, et al.
Published: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)
by: Ko, Hyunwoo, et al.
Published: (2025)
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation
by: Aleem, Sidra, et al.
Published: (2024)
by: Aleem, Sidra, et al.
Published: (2024)
A Monte Carlo Language Model Pipeline for Zero-Shot Sociopolitical Event Extraction
by: Cai, Erica, et al.
Published: (2023)
by: Cai, Erica, et al.
Published: (2023)
Spatial Audio Motion Understanding and Reasoning
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
by: Zhang, Yingji, et al.
Published: (2026)
by: Zhang, Yingji, et al.
Published: (2026)
Coordinates from Context: Using LLMs to Ground Complex Location References
by: Masis, Tessa, et al.
Published: (2025)
by: Masis, Tessa, et al.
Published: (2025)
Where on Earth Do Users Say They Are?: Geo-Entity Linking for Noisy Multilingual User Input
by: Masis, Tessa, et al.
Published: (2024)
by: Masis, Tessa, et al.
Published: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
by: Russell, Sam O'Connor, et al.
Published: (2025)
by: Russell, Sam O'Connor, et al.
Published: (2025)
Latin Treebanks in Review: An Evaluation of Morphological Tagging Across Time
by: Hudspeth, Marisa, et al.
Published: (2024)
by: Hudspeth, Marisa, et al.
Published: (2024)
Evaluating Morphological Alignment of Tokenizers in 70 Languages
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
FSLI: An Interpretable Formal Semantic System for One-Dimensional Ordering Inference
by: Alkhairy, Maha, et al.
Published: (2025)
by: Alkhairy, Maha, et al.
Published: (2025)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
How Can Large Language Models Understand Spatial-Temporal Data?
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
by: Li, Xingxuan, et al.
Published: (2024)
by: Li, Xingxuan, et al.
Published: (2024)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures
by: Sidhu, Risham, et al.
Published: (2026)
by: Sidhu, Risham, et al.
Published: (2026)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
by: Subbiah, Melanie, et al.
Published: (2025)
by: Subbiah, Melanie, et al.
Published: (2025)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering
by: Hudspeth, Marisa, et al.
Published: (2026)
by: Hudspeth, Marisa, et al.
Published: (2026)
Contextual morphologically-guided tokenization for Latin encoder models
by: Hudspeth, Marisa, et al.
Published: (2025)
by: Hudspeth, Marisa, et al.
Published: (2025)
Can Large Language Models Create New Knowledge for Spatial Reasoning Tasks?
by: Greatrix, Thomas, et al.
Published: (2024)
by: Greatrix, Thomas, et al.
Published: (2024)
Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?
by: Li, Jason, et al.
Published: (2025)
by: Li, Jason, et al.
Published: (2025)
Can Large Language Models Understand Context?
by: Zhu, Yilun, et al.
Published: (2024)
by: Zhu, Yilun, et al.
Published: (2024)
TextBandit: Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks
by: Lim, Jimin, et al.
Published: (2025)
by: Lim, Jimin, et al.
Published: (2025)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
by: Du, Mengfei, et al.
Published: (2024)
by: Du, Mengfei, et al.
Published: (2024)
Understanding and Patching Compositional Reasoning in LLMs
by: Li, Zhaoyi, et al.
Published: (2024)
by: Li, Zhaoyi, et al.
Published: (2024)
Similar Items
-
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026) -
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026) -
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026) -
Metamorphic Testing for Pose Estimation Systems
by: Duran, Matias, et al.
Published: (2025) -
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024)