Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xingrui, Ma, Wufei, Zhang, Tiezheng, de Melo, Celso M, Chen, Jieneng, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
Pursuing Minimal Sufficiency in Spatial Reasoning
von: Guo, Yejie, et al.
Veröffentlicht: (2025)
von: Guo, Yejie, et al.
Veröffentlicht: (2025)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
von: Ma, Wenxin, et al.
Veröffentlicht: (2026)
von: Ma, Wenxin, et al.
Veröffentlicht: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
von: Guan, Yaohan, et al.
Veröffentlicht: (2026)
von: Guan, Yaohan, et al.
Veröffentlicht: (2026)
Leveraging AI Predicted and Expert Revised Annotations in Interactive Segmentation: Continual Tuning or Full Training?
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2024)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2026)
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2026)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
von: Zhong, Shanshan, et al.
Veröffentlicht: (2025)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2025)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
von: Mou, Tingshu, et al.
Veröffentlicht: (2026)
von: Mou, Tingshu, et al.
Veröffentlicht: (2026)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
von: Zhao, Bokai, et al.
Veröffentlicht: (2026)
von: Zhao, Bokai, et al.
Veröffentlicht: (2026)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
von: Wang, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyan, et al.
Veröffentlicht: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
von: Liu, Chonghan, et al.
Veröffentlicht: (2025)
von: Liu, Chonghan, et al.
Veröffentlicht: (2025)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
von: Sharma, Arun
Veröffentlicht: (2026)
von: Sharma, Arun
Veröffentlicht: (2026)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
von: Liu, Zishan, et al.
Veröffentlicht: (2026)
von: Liu, Zishan, et al.
Veröffentlicht: (2026)
Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
von: Frisoni, Giacomo, et al.
Veröffentlicht: (2026)
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
von: Wu, Nemin, et al.
Veröffentlicht: (2024)
von: Wu, Nemin, et al.
Veröffentlicht: (2024)
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2026)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2026)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
von: Ma, Wufei, et al.
Veröffentlicht: (2025) -
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
von: Ma, Wufei, et al.
Veröffentlicht: (2024) -
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
von: Ma, Wufei, et al.
Veröffentlicht: (2025) -
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
von: Wang, Xingrui, et al.
Veröffentlicht: (2024) -
Pursuing Minimal Sufficiency in Spatial Reasoning
von: Guo, Yejie, et al.
Veröffentlicht: (2025)