Towards Understanding Graphical Perception in Large Multimodal Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Kai, Yang, Jianwei, Inala, Jeevana Priya, Singh, Chandan, Gao, Jianfeng, Su, Yu, Wang, Chenglong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916652374818816
author Zhang, Kai
Yang, Jianwei
Inala, Jeevana Priya
Singh, Chandan
Gao, Jianfeng
Su, Yu
Wang, Chenglong
author_facet Zhang, Kai
Yang, Jianwei
Inala, Jeevana Priya
Singh, Chandan
Gao, Jianfeng
Su, Yu
Wang, Chenglong
contents Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprisingly found that these models struggle with simple tasks on infographics that require perception only. As existing benchmarks primarily focus on end tasks that require various abilities, they provide limited, fine-grained insights into the limitations of the models' perception abilities. To address this gap, we leverage the theory of graphical perception, an approach used to study how humans decode visual information encoded on charts and graphs, to develop an evaluation framework for analyzing gaps in LMMs' perception abilities in charts. With automated task generation and response evaluation designs, our framework enables comprehensive and controlled testing of LMMs' graphical perception across diverse chart types, visual elements, and task types. We apply our framework to evaluate and diagnose the perception capabilities of state-of-the-art LMMs at three granularity levels (chart, visual element, and pixel). Our findings underscore several critical limitations of current state-of-the-art LMMs, including GPT-4o: their inability to (1) generalize across chart types, (2) understand fundamental visual elements, and (3) cross reference values within a chart. These insights provide guidance for future improvements in perception abilities of LMMs. The evaluation framework and labeled data are publicly available at https://github.com/microsoft/lmm-graphical-perception.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10857
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Understanding Graphical Perception in Large Multimodal Models
Zhang, Kai
Yang, Jianwei
Inala, Jeevana Priya
Singh, Chandan
Gao, Jianfeng
Su, Yu
Wang, Chenglong
Graphics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprisingly found that these models struggle with simple tasks on infographics that require perception only. As existing benchmarks primarily focus on end tasks that require various abilities, they provide limited, fine-grained insights into the limitations of the models' perception abilities. To address this gap, we leverage the theory of graphical perception, an approach used to study how humans decode visual information encoded on charts and graphs, to develop an evaluation framework for analyzing gaps in LMMs' perception abilities in charts. With automated task generation and response evaluation designs, our framework enables comprehensive and controlled testing of LMMs' graphical perception across diverse chart types, visual elements, and task types. We apply our framework to evaluate and diagnose the perception capabilities of state-of-the-art LMMs at three granularity levels (chart, visual element, and pixel). Our findings underscore several critical limitations of current state-of-the-art LMMs, including GPT-4o: their inability to (1) generalize across chart types, (2) understand fundamental visual elements, and (3) cross reference values within a chart. These insights provide guidance for future improvements in perception abilities of LMMs. The evaluation framework and labeled data are publicly available at https://github.com/microsoft/lmm-graphical-perception.
title Towards Understanding Graphical Perception in Large Multimodal Models
topic Graphics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10857