BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Junyan, Jiang, Dongzhi, He, Jun, Zhou, Baichuan, Huang, Zilong, Yan, Zhiyuan, Li, Hongsheng, He, Conghui, Li, Weijia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912640940376064
author Ye, Junyan
Jiang, Dongzhi
He, Jun
Zhou, Baichuan
Huang, Zilong
Yan, Zhiyuan
Li, Hongsheng
He, Conghui
Li, Weijia
author_facet Ye, Junyan
Jiang, Dongzhi
He, Jun
Zhou, Baichuan
Huang, Zilong
Yan, Zhiyuan
Li, Hongsheng
He, Conghui
Li, Weijia
contents Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating visual input as replaceable context. To address this gap, we introduce BLINK-Twice, a vision-centric reasoning benchmark grounded in challenging perceptual tasks. Instead of relying on external knowledge, our tasks require models to reason from visual content alone, shifting the focus from language-based to image-grounded reasoning. Compared to prior perception benchmarks, it moves beyond shallow perception ("see") and requires fine-grained observation and analytical reasoning ("observe"). BLINK-Twice integrates three core components: seven types of visual challenges for testing visual reasoning, natural adversarial image pairs that enforce reliance on visual content, and annotated reasoning chains for fine-grained evaluation of the reasoning process rather than final answers alone. We evaluate 20 leading MLLMs, including 12 foundation models and 8 reasoning-enhanced models. BLINK-Twice poses a significant challenge to current models. While existing reasoning strategies in the language space-such as chain-of-thought or self-criticism can improve performance, they often result in unstable and redundant reasoning. We observe that repeated image observation improves performance across models, and active visual interaction, as demonstrated by models like o3, highlights the need for a new paradigm for vision reasoning. The dataset is publicly available at https://github.com/PicoTrex/BLINK-Twice
format Preprint
id arxiv_https___arxiv_org_abs_2510_09361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
Ye, Junyan
Jiang, Dongzhi
He, Jun
Zhou, Baichuan
Huang, Zilong
Yan, Zhiyuan
Li, Hongsheng
He, Conghui
Li, Weijia
Computer Vision and Pattern Recognition
Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating visual input as replaceable context. To address this gap, we introduce BLINK-Twice, a vision-centric reasoning benchmark grounded in challenging perceptual tasks. Instead of relying on external knowledge, our tasks require models to reason from visual content alone, shifting the focus from language-based to image-grounded reasoning. Compared to prior perception benchmarks, it moves beyond shallow perception ("see") and requires fine-grained observation and analytical reasoning ("observe"). BLINK-Twice integrates three core components: seven types of visual challenges for testing visual reasoning, natural adversarial image pairs that enforce reliance on visual content, and annotated reasoning chains for fine-grained evaluation of the reasoning process rather than final answers alone. We evaluate 20 leading MLLMs, including 12 foundation models and 8 reasoning-enhanced models. BLINK-Twice poses a significant challenge to current models. While existing reasoning strategies in the language space-such as chain-of-thought or self-criticism can improve performance, they often result in unstable and redundant reasoning. We observe that repeated image observation improves performance across models, and active visual interaction, as demonstrated by models like o3, highlights the need for a new paradigm for vision reasoning. The dataset is publicly available at https://github.com/PicoTrex/BLINK-Twice
title BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.09361