Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Zhikai, Sun, Jiashuo, Zhang, Wenqi, Hu, Zhiqiang, Li, Xin, Wang, Fan, Zhao, Deli
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910940434268160
author Wang, Zhikai
Sun, Jiashuo
Zhang, Wenqi
Hu, Zhiqiang
Li, Xin
Wang, Fan
Zhao, Deli
author_facet Wang, Zhikai
Sun, Jiashuo
Zhang, Wenqi
Hu, Zhiqiang
Li, Xin
Wang, Fan
Zhao, Deli
contents Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual question answering. However, current benchmarks typically focus on knowledge-centric evaluations that assess domain-specific expertise, often neglecting the core ability to reason about fundamental mathematical elements and visual concepts. We identify a gap in evaluating elementary-level math problems, which rely on explicit visual dependencies-requiring models to discern, integrate, and reason across multiple images while incorporating commonsense knowledge, all of which are crucial for advancing toward broader AGI capabilities. To address this gap, we introduce VCBENCH, a comprehensive benchmark for multimodal mathematical reasoning with explicit visual dependencies. VCBENCH includes 1,720 problems across six cognitive domains, featuring 6,697 images (averaging 3.9 per question) to ensure multi-image reasoning. We evaluate 26 state-of-the-art LVLMs on VCBENCH, revealing substantial performance disparities, with even the top models unable to exceed 50% accuracy. Our findings highlight the ongoing challenges in visual-mathematical integration and suggest avenues for future LVLM advancements. The project can be found at https://alibaba-damo-academy.github.io/VCBench/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18589
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
Wang, Zhikai
Sun, Jiashuo
Zhang, Wenqi
Hu, Zhiqiang
Li, Xin
Wang, Fan
Zhao, Deli
Computer Vision and Pattern Recognition
Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual question answering. However, current benchmarks typically focus on knowledge-centric evaluations that assess domain-specific expertise, often neglecting the core ability to reason about fundamental mathematical elements and visual concepts. We identify a gap in evaluating elementary-level math problems, which rely on explicit visual dependencies-requiring models to discern, integrate, and reason across multiple images while incorporating commonsense knowledge, all of which are crucial for advancing toward broader AGI capabilities. To address this gap, we introduce VCBENCH, a comprehensive benchmark for multimodal mathematical reasoning with explicit visual dependencies. VCBENCH includes 1,720 problems across six cognitive domains, featuring 6,697 images (averaging 3.9 per question) to ensure multi-image reasoning. We evaluate 26 state-of-the-art LVLMs on VCBENCH, revealing substantial performance disparities, with even the top models unable to exceed 50% accuracy. Our findings highlight the ongoing challenges in visual-mathematical integration and suggest avenues for future LVLM advancements. The project can be found at https://alibaba-damo-academy.github.io/VCBench/.
title Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.18589