Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Linjie, Bigverdi, Mahtab, Gu, Jiawei, Ma, Zixian, Yang, Yinuo, Li, Ziang, Choi, Yejin, Krishna, Ranjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910989206683648
author Li, Linjie
Bigverdi, Mahtab
Gu, Jiawei
Ma, Zixian
Yang, Yinuo
Li, Ziang
Choi, Yejin
Krishna, Ranjay
author_facet Li, Linjie
Bigverdi, Mahtab
Gu, Jiawei
Ma, Zixian
Yang, Yinuo
Li, Ziang
Choi, Yejin
Krishna, Ranjay
contents Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the complexities of non-verbal, multi-step visual simulation. We introduce STARE(Spatial Transformations and Reasoning Evaluation), a benchmark designed to rigorously evaluate multimodal large language models on tasks better solved through multi-step visual simulation. STARE features 4K tasks spanning foundational geometric transformations (2D and 3D), integrated spatial reasoning (cube net folding and tangram puzzles), and real-world spatial reasoning (perspective and temporal reasoning), reflecting practical cognitive challenges like object assembly, mechanical diagram interpretation, and everyday spatial navigation. Our evaluations show that models excel at reasoning over simpler 2D transformations, but perform close to random chance on more complex tasks like 3D cube net folding and tangram puzzles that require multi-step visual simulations. Humans achieve near-perfect accuracy but take considerable time (up to 28.9s) on complex tasks, significantly speeding up (down by 7.5 seconds on average) with intermediate visual simulations. In contrast, models exhibit inconsistent performance gains from visual simulations, improving on most tasks but declining in specific cases like tangram puzzles (GPT-4o, o1) and cube net folding (Claude-3.5, Gemini-2.0 Flash), indicating that models may not know how to effectively leverage intermediate visual information.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04633
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
Li, Linjie
Bigverdi, Mahtab
Gu, Jiawei
Ma, Zixian
Yang, Yinuo
Li, Ziang
Choi, Yejin
Krishna, Ranjay
Computer Vision and Pattern Recognition
Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the complexities of non-verbal, multi-step visual simulation. We introduce STARE(Spatial Transformations and Reasoning Evaluation), a benchmark designed to rigorously evaluate multimodal large language models on tasks better solved through multi-step visual simulation. STARE features 4K tasks spanning foundational geometric transformations (2D and 3D), integrated spatial reasoning (cube net folding and tangram puzzles), and real-world spatial reasoning (perspective and temporal reasoning), reflecting practical cognitive challenges like object assembly, mechanical diagram interpretation, and everyday spatial navigation. Our evaluations show that models excel at reasoning over simpler 2D transformations, but perform close to random chance on more complex tasks like 3D cube net folding and tangram puzzles that require multi-step visual simulations. Humans achieve near-perfect accuracy but take considerable time (up to 28.9s) on complex tasks, significantly speeding up (down by 7.5 seconds on average) with intermediate visual simulations. In contrast, models exhibit inconsistent performance gains from visual simulations, improving on most tasks but declining in specific cases like tangram puzzles (GPT-4o, o1) and cube net folding (Claude-3.5, Gemini-2.0 Flash), indicating that models may not know how to effectively leverage intermediate visual information.
title Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04633