MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeller, Jana, Wiedemer, Thaddäus, Li, Fanfei, Klein, Thomas, Mayilvahanan, Prasanna, Bethge, Matthias, Wichmann, Felix, Cotterell, Ryan, Brendel, Wieland
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914301817651200
author Zeller, Jana
Wiedemer, Thaddäus
Li, Fanfei
Klein, Thomas
Mayilvahanan, Prasanna
Bethge, Matthias
Wichmann, Felix
Cotterell, Ryan
Brendel, Wieland
author_facet Zeller, Jana
Wiedemer, Thaddäus
Li, Fanfei
Klein, Thomas
Mayilvahanan, Prasanna
Bethge, Matthias
Wichmann, Felix
Cotterell, Ryan
Brendel, Wieland
contents Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human mental imagery. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, visual thoughts do not yet benefit model reasoning. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02465
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
Zeller, Jana
Wiedemer, Thaddäus
Li, Fanfei
Klein, Thomas
Mayilvahanan, Prasanna
Bethge, Matthias
Wichmann, Felix
Cotterell, Ryan
Brendel, Wieland
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human mental imagery. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, visual thoughts do not yet benefit model reasoning. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.
title MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.02465