Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sinha, Rohit, Kanade, Aditya, Kancheti, Sai Srinivas, Balasubramanian, Vineeth N, Ganu, Tanuja
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917416067399680
author Sinha, Rohit
Kanade, Aditya
Kancheti, Sai Srinivas
Balasubramanian, Vineeth N
Ganu, Tanuja
author_facet Sinha, Rohit
Kanade, Aditya
Kancheti, Sai Srinivas
Balasubramanian, Vineeth N
Ganu, Tanuja
contents Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel "A-R-T" taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16054
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
Sinha, Rohit
Kanade, Aditya
Kancheti, Sai Srinivas
Balasubramanian, Vineeth N
Ganu, Tanuja
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel "A-R-T" taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks.
title Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.16054