iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mayer, Julius, Ballout, Mohamad, Jassim, Serwan, Nezami, Farbod Nosrat, Bruni, Elia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908567599054848
author Mayer, Julius
Ballout, Mohamad
Jassim, Serwan
Nezami, Farbod Nosrat
Bruni, Elia
author_facet Mayer, Julius
Ballout, Mohamad
Jassim, Serwan
Nezami, Farbod Nosrat
Bruni, Elia
contents Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. \mbox{iVISPAR} is based on a variant of the sliding tile puzzle, a classic problem that demands logical planning, spatial awareness, and multi-step reasoning. The benchmark supports visual 3D, 2D, and text-based input modalities, enabling comprehensive assessments of VLMs' planning and reasoning skills. We evaluate a broad suite of state-of-the-art open-source and closed-source VLMs, comparing their performance while also providing optimal path solutions and a human baseline to assess the task's complexity and feasibility for humans. Results indicate that while VLMs perform better on 2D tasks compared to 3D or text-based settings, they struggle with complex spatial configurations and consistently fall short of human performance, illustrating the persistent challenge of visual alignment. This underscores critical gaps in current VLM capabilities, highlighting their limitations in achieving human-level cognition. Project website: https://microcosm.ai/ivispar
format Preprint
id arxiv_https___arxiv_org_abs_2502_03214
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
Mayer, Julius
Ballout, Mohamad
Jassim, Serwan
Nezami, Farbod Nosrat
Bruni, Elia
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. \mbox{iVISPAR} is based on a variant of the sliding tile puzzle, a classic problem that demands logical planning, spatial awareness, and multi-step reasoning. The benchmark supports visual 3D, 2D, and text-based input modalities, enabling comprehensive assessments of VLMs' planning and reasoning skills. We evaluate a broad suite of state-of-the-art open-source and closed-source VLMs, comparing their performance while also providing optimal path solutions and a human baseline to assess the task's complexity and feasibility for humans. Results indicate that while VLMs perform better on 2D tasks compared to 3D or text-based settings, they struggle with complex spatial configurations and consistently fall short of human performance, illustrating the persistent challenge of visual alignment. This underscores critical gaps in current VLM capabilities, highlighting their limitations in achieving human-level cognition. Project website: https://microcosm.ai/ivispar
title iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.03214