RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Agarwal, Amit, Patel, Hitesh Laxmichand, Panda, Srikant, Meghwani, Hansa, Singh, Jyotika, Dua, Karan, Li, Paul, Sheng, Tao, Ravi, Sujith, Roth, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909813004304384
author Agarwal, Amit
Patel, Hitesh Laxmichand
Panda, Srikant
Meghwani, Hansa
Singh, Jyotika
Dua, Karan
Li, Paul
Sheng, Tao
Ravi, Sujith
Roth, Dan
author_facet Agarwal, Amit
Patel, Hitesh Laxmichand
Panda, Srikant
Meghwani, Hansa
Singh, Jyotika
Dua, Karan
Li, Paul
Sheng, Tao
Ravi, Sujith
Roth, Dan
contents Multimodal Large Language Models (MLLMs) have achieved impressive results on vision-language benchmarks, yet it remains unclear whether these benchmarks assess genuine global reasoning or allow success via localized visual cues. Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. We introduce Region Comprehension Index (RCI), the first model-based score to directly quantify a dataset's reliance on global versus local visual information. RCI systematically compares reference-model performance on image patches versus full images, revealing if tasks require holistic image understanding or can be solved with partial or localized visual cues. When applying RCI to 13 widely used multimodal benchmarks, we observed that most of them favor localized reasoning and exhibit significant spatial biases, indicating potential risks in real-world applications. RCI equips researchers & practitioners with an actionable tool for diagnosing & mitigating these biases, enabling the construction of datasets and benchmarks to foster the development of robust, enterprise-ready multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23673
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
Agarwal, Amit
Patel, Hitesh Laxmichand
Panda, Srikant
Meghwani, Hansa
Singh, Jyotika
Dua, Karan
Li, Paul
Sheng, Tao
Ravi, Sujith
Roth, Dan
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
68T45, 68T50
I.2.7; I.2.10; I.4.7; I.4.8
Multimodal Large Language Models (MLLMs) have achieved impressive results on vision-language benchmarks, yet it remains unclear whether these benchmarks assess genuine global reasoning or allow success via localized visual cues. Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. We introduce Region Comprehension Index (RCI), the first model-based score to directly quantify a dataset's reliance on global versus local visual information. RCI systematically compares reference-model performance on image patches versus full images, revealing if tasks require holistic image understanding or can be solved with partial or localized visual cues. When applying RCI to 13 widely used multimodal benchmarks, we observed that most of them favor localized reasoning and exhibit significant spatial biases, indicating potential risks in real-world applications. RCI equips researchers & practitioners with an actionable tool for diagnosing & mitigating these biases, enabling the construction of datasets and benchmarks to foster the development of robust, enterprise-ready multimodal systems.
title RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
68T45, 68T50
I.2.7; I.2.10; I.4.7; I.4.8
url https://arxiv.org/abs/2509.23673