When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Hongcheng, Wang, Pingjie, Wang, Yuhao, Ou, Siqu, Wang, Yanfeng, Wang, Yu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918162477350912
author Liu, Hongcheng
Wang, Pingjie
Wang, Yuhao
Ou, Siqu
Wang, Yanfeng
Wang, Yu
author_facet Liu, Hongcheng
Wang, Pingjie
Wang, Yuhao
Ou, Siqu
Wang, Yanfeng
Wang, Yu
contents Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete information. This setup is misaligned with real-world use, where seeing is not enough. This raises a fundamental question: Can MLLMs actively acquire missing evidence under incomplete information? To bridge this gap, we require the MLLMs to actively acquire missing evidence and iteratively refine decisions under incomplete information, by selecting a target image from a candidate pool without task-specific priors. To support systematic study, we propose GuessBench, a benchmark with both perception-oriented and knowledge-oriented images for evaluating active reasoning in MLLMs. We evaluate 20 superior MLLMs and find that performance on active reasoning lags far behind it on passive settings, indicating substantial room for improvement. Further analysis identifies fine-grained perception and timely decision-making as key challenges. Ablation studies show that perceptual enhancements benefit smaller models, whereas thinking-oriented methods provide consistent gains across model sizes. These results suggest promising directions for future research on multimodal active reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
Liu, Hongcheng
Wang, Pingjie
Wang, Yuhao
Ou, Siqu
Wang, Yanfeng
Wang, Yu
Computation and Language
Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete information. This setup is misaligned with real-world use, where seeing is not enough. This raises a fundamental question: Can MLLMs actively acquire missing evidence under incomplete information? To bridge this gap, we require the MLLMs to actively acquire missing evidence and iteratively refine decisions under incomplete information, by selecting a target image from a candidate pool without task-specific priors. To support systematic study, we propose GuessBench, a benchmark with both perception-oriented and knowledge-oriented images for evaluating active reasoning in MLLMs. We evaluate 20 superior MLLMs and find that performance on active reasoning lags far behind it on passive settings, indicating substantial room for improvement. Further analysis identifies fine-grained perception and timely decision-making as key challenges. Ablation studies show that perceptual enhancements benefit smaller models, whereas thinking-oriented methods provide consistent gains across model sizes. These results suggest promising directions for future research on multimodal active reasoning.
title When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
topic Computation and Language
url https://arxiv.org/abs/2510.15421