SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Siting, Pei, Minnan, Sun, Luoyang, Deng, Cheng, Li, Yuchen, Shao, Kun, Tian, Zheng, Zhang, Haifeng, Wang, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911516656140288
author Wang, Siting
Pei, Minnan
Sun, Luoyang
Deng, Cheng
Li, Yuchen
Shao, Kun
Tian, Zheng
Zhang, Haifeng
Wang, Jun
author_facet Wang, Siting
Pei, Minnan
Sun, Luoyang
Deng, Cheng
Li, Yuchen
Shao, Kun
Tian, Zheng
Zhang, Haifeng
Wang, Jun
contents Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated as a spatial skill. This reliance on publicly sourced problems from IQ tests or math competitions risks data contamination and compromises assessment reliability. To this end, we introduce SpatialViz-Bench, a comprehensive multi-modal benchmark for spatial visualization with 12 tasks across 4 sub-abilities, comprising 1,180 programmatically generated problems, a scalable framework that allows for expansion to ensure fair and continuously reliable evaluations. Our evaluation of 27 Multi-modal Large Language Models (MLLMs) reveals wide performance variations, demonstrates the benchmark's strong discriminative power, and uncovers counter-intuitive findings: Chain-of-Thought (CoT) prompting paradoxically degrades accuracy on open-source models. Through statistical and qualitative analysis of error types, SpatialViz-Bench demonstrates that state-of-the-art MLLMs exhibit deficiencies in spatial visualization tasks, thereby addressing a significant lacuna in the field. The benchmark data and evaluation code are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
Wang, Siting
Pei, Minnan
Sun, Luoyang
Deng, Cheng
Li, Yuchen
Shao, Kun
Tian, Zheng
Zhang, Haifeng
Wang, Jun
Computer Vision and Pattern Recognition
Computation and Language
Human-Computer Interaction
Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated as a spatial skill. This reliance on publicly sourced problems from IQ tests or math competitions risks data contamination and compromises assessment reliability. To this end, we introduce SpatialViz-Bench, a comprehensive multi-modal benchmark for spatial visualization with 12 tasks across 4 sub-abilities, comprising 1,180 programmatically generated problems, a scalable framework that allows for expansion to ensure fair and continuously reliable evaluations. Our evaluation of 27 Multi-modal Large Language Models (MLLMs) reveals wide performance variations, demonstrates the benchmark's strong discriminative power, and uncovers counter-intuitive findings: Chain-of-Thought (CoT) prompting paradoxically degrades accuracy on open-source models. Through statistical and qualitative analysis of error types, SpatialViz-Bench demonstrates that state-of-the-art MLLMs exhibit deficiencies in spatial visualization tasks, thereby addressing a significant lacuna in the field. The benchmark data and evaluation code are publicly available.
title SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
topic Computer Vision and Pattern Recognition
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2507.07610