InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Tianchi, Lin, Minzhi, Liu, Mengchen, Ye, Yilin, Chen, Changjian, Liu, Shixia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918176731693056
author Xie, Tianchi
Lin, Minzhi
Liu, Mengchen
Ye, Yilin
Chen, Changjian
Liu, Shixia
author_facet Xie, Tianchi
Lin, Minzhi
Liu, Mengchen
Ye, Yilin
Chen, Changjian
Liu, Shixia
contents Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, existing visual-question answering benchmarks fall short in evaluating these capabilities of MLLMs due to the lack of paired plain charts and visual-element-based questions. To bridge this gap, we introduce InfoChartQA, a benchmark for evaluating MLLMs on infographic chart understanding. It includes 5,642 pairs of infographic and plain charts, each sharing the same underlying data but differing in visual presentations. We further design visual-element-based questions to capture their unique visual designs and communicative intent. Evaluation of 20 MLLMs reveals a substantial performance decline on infographic charts, particularly for visual-element-based questions related to metaphors. The paired infographic and plain charts enable fine-grained error analysis and ablation studies, which highlight new opportunities for advancing MLLMs in infographic chart understanding. We release InfoChartQA at https://github.com/CoolDawnAnt/InfoChartQA.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19028
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
Xie, Tianchi
Lin, Minzhi
Liu, Mengchen
Ye, Yilin
Chen, Changjian
Liu, Shixia
Computer Vision and Pattern Recognition
Artificial Intelligence
Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, existing visual-question answering benchmarks fall short in evaluating these capabilities of MLLMs due to the lack of paired plain charts and visual-element-based questions. To bridge this gap, we introduce InfoChartQA, a benchmark for evaluating MLLMs on infographic chart understanding. It includes 5,642 pairs of infographic and plain charts, each sharing the same underlying data but differing in visual presentations. We further design visual-element-based questions to capture their unique visual designs and communicative intent. Evaluation of 20 MLLMs reveals a substantial performance decline on infographic charts, particularly for visual-element-based questions related to metaphors. The paired infographic and plain charts enable fine-grained error analysis and ablation studies, which highlight new opportunities for advancing MLLMs in infographic chart understanding. We release InfoChartQA at https://github.com/CoolDawnAnt/InfoChartQA.
title InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.19028