VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Jing, Guo, Junjia, Liang, Susan, Sun, Guangyu, Song, Luchuan, Tang, Yunlong, He, Jinxi, Wu, Jiarui, Vosoughi, Ali, Chen, Chen, Xu, Chenliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917957262639104
author Bi, Jing
Guo, Junjia
Liang, Susan
Sun, Guangyu
Song, Luchuan
Tang, Yunlong
He, Jinxi
Wu, Jiarui
Vosoughi, Ali
Chen, Chen
Xu, Chenliang
author_facet Bi, Jing
Guo, Junjia
Liang, Susan
Sun, Guangyu
Song, Luchuan
Tang, Yunlong
He, Jinxi
Wu, Jiarui
Vosoughi, Ali
Chen, Chen
Xu, Chenliang
contents Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across language and vision-language tasks, existing benchmarks primarily measure recognition-based skills and inadequately assess true visual reasoning capabilities. To bridge this critical gap, we introduce VERIFY, a benchmark explicitly designed to isolate and rigorously evaluate the visual reasoning capabilities of state-of-the-art MLLMs. VERIFY compels models to reason primarily from visual information, providing minimal textual context to reduce reliance on domain-specific knowledge and linguistic biases. Each problem is accompanied by a human-annotated reasoning path, making it the first to provide in-depth evaluation of model decision-making processes. Additionally, we propose novel metrics that assess visual reasoning fidelity beyond mere accuracy, highlighting critical imbalances in current model reasoning patterns. Our comprehensive benchmarking of leading MLLMs uncovers significant limitations, underscoring the need for a balanced and holistic approach to both perception and reasoning. For more teaser and testing, visit our project page (https://verify-eqh.pages.dev/).
format Preprint
id arxiv_https___arxiv_org_abs_2503_11557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
Bi, Jing
Guo, Junjia
Liang, Susan
Sun, Guangyu
Song, Luchuan
Tang, Yunlong
He, Jinxi
Wu, Jiarui
Vosoughi, Ali
Chen, Chen
Xu, Chenliang
Computer Vision and Pattern Recognition
Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across language and vision-language tasks, existing benchmarks primarily measure recognition-based skills and inadequately assess true visual reasoning capabilities. To bridge this critical gap, we introduce VERIFY, a benchmark explicitly designed to isolate and rigorously evaluate the visual reasoning capabilities of state-of-the-art MLLMs. VERIFY compels models to reason primarily from visual information, providing minimal textual context to reduce reliance on domain-specific knowledge and linguistic biases. Each problem is accompanied by a human-annotated reasoning path, making it the first to provide in-depth evaluation of model decision-making processes. Additionally, we propose novel metrics that assess visual reasoning fidelity beyond mere accuracy, highlighting critical imbalances in current model reasoning patterns. Our comprehensive benchmarking of leading MLLMs uncovers significant limitations, underscoring the need for a balanced and holistic approach to both perception and reasoning. For more teaser and testing, visit our project page (https://verify-eqh.pages.dev/).
title VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.11557