AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xintong, Zhang, Xiaowen, Wu, Jingrong, Gao, Zhi, Yan, Shilin, Diao, Zhenxin, Gao, Kunpeng, Chen, Xuanyan, Wu, Yuwei, Jia, Yunde, Li, Qing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908947043057664
author Zhang, Xintong
Zhang, Xiaowen
Wu, Jingrong
Gao, Zhi
Yan, Shilin
Diao, Zhenxin
Gao, Kunpeng
Chen, Xuanyan
Wu, Yuwei
Jia, Yunde
Li, Qing
author_facet Zhang, Xintong
Zhang, Xiaowen
Wu, Jingrong
Gao, Zhi
Yan, Shilin
Diao, Zhenxin
Gao, Kunpeng
Chen, Xuanyan
Wu, Yuwei
Jia, Yunde
Li, Qing
contents Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text reasoning to enhance both effectiveness and efficiency. However, existing evaluations rely on static difficulty labels and simplistic metrics, which fail to capture the dynamic nature of difficulty relative to varying model capacities. Consequently, they obscure the distinction between adaptive mode selection and general performance while neglecting fine-grained process analyses. In this paper, we propose AdaptMMBench, a comprehensive benchmark for adaptive multimodal reasoning across five domains: real-world, OCR, GUI, knowledge, and math, encompassing both direct perception and complex reasoning tasks. AdaptMMBench utilizes a Matthews Correlation Coefficient (MCC) metric to evaluate the selection rationality of different reasoning modes, isolating this meta-cognition ability by dynamically identifying task difficulties based on models' capability boundaries. Moreover, AdaptMMBench facilitates multi-dimensional process evaluation across key step coverage, tool effectiveness, and computational efficiency. Our evaluation reveals that while adaptive mode selection scales with model capacity, it notably decouples from final accuracy. Conversely, key step coverage aligns with performance, though tool effectiveness remains highly inconsistent across model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
Zhang, Xintong
Zhang, Xiaowen
Wu, Jingrong
Gao, Zhi
Yan, Shilin
Diao, Zhenxin
Gao, Kunpeng
Chen, Xuanyan
Wu, Yuwei
Jia, Yunde
Li, Qing
Computer Vision and Pattern Recognition
Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text reasoning to enhance both effectiveness and efficiency. However, existing evaluations rely on static difficulty labels and simplistic metrics, which fail to capture the dynamic nature of difficulty relative to varying model capacities. Consequently, they obscure the distinction between adaptive mode selection and general performance while neglecting fine-grained process analyses. In this paper, we propose AdaptMMBench, a comprehensive benchmark for adaptive multimodal reasoning across five domains: real-world, OCR, GUI, knowledge, and math, encompassing both direct perception and complex reasoning tasks. AdaptMMBench utilizes a Matthews Correlation Coefficient (MCC) metric to evaluate the selection rationality of different reasoning modes, isolating this meta-cognition ability by dynamically identifying task difficulties based on models' capability boundaries. Moreover, AdaptMMBench facilitates multi-dimensional process evaluation across key step coverage, tool effectiveness, and computational efficiency. Our evaluation reveals that while adaptive mode selection scales with model capacity, it notably decouples from final accuracy. Conversely, key step coverage aligns with performance, though tool effectiveness remains highly inconsistent across model architectures.
title AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02676