CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Zirui, Xia, Mengzhou, He, Luxi, Chen, Howard, Liu, Yitao, Zhu, Richard, Liang, Kaiqu, Wu, Xindi, Liu, Haotian, Malladi, Sadhika, Chevalier, Alexis, Arora, Sanjeev, Chen, Danqi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929400735334400
author Wang, Zirui
Xia, Mengzhou
He, Luxi
Chen, Howard
Liu, Yitao
Zhu, Richard
Liang, Kaiqu
Wu, Xindi
Liu, Haotian
Malladi, Sadhika
Chevalier, Alexis
Arora, Sanjeev
Chen, Danqi
author_facet Wang, Zirui
Xia, Mengzhou
He, Luxi
Chen, Howard
Liu, Yitao
Zhu, Richard
Liang, Kaiqu
Wu, Xindi
Liu, Haotian
Malladi, Sadhika
Chevalier, Alexis
Arora, Sanjeev
Chen, Danqi
contents Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2406_18521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Wang, Zirui
Xia, Mengzhou
He, Luxi
Chen, Howard
Liu, Yitao
Zhu, Richard
Liang, Kaiqu
Wu, Xindi
Liu, Haotian
Malladi, Sadhika
Chevalier, Alexis
Arora, Sanjeev
Chen, Danqi
Computation and Language
Computer Vision and Pattern Recognition
Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/
title CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.18521