POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yichen, Chen, Liangyu, Zhang, Liang, Ma, Jianzhe, Wang, Wenxuan, Jin, Qin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911361147076608
author Xu, Yichen
Chen, Liangyu
Zhang, Liang
Ma, Jianzhe
Wang, Wenxuan
Jin, Qin
author_facet Xu, Yichen
Chen, Liangyu
Zhang, Liang
Ma, Jianzhe
Wang, Wenxuan
Jin, Qin
contents Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. To address this limitation, we introduce PolyChartQA, the first large-scale multilingual benchmark for chart question answering, comprising 22,606 charts and 26,151 QA pairs across 10 diverse languages. PolyChartQA is constructed through a scalable pipeline that enables efficient multilingual chart generation via data translation and code reuse, supported by LLM-based translation and rigorous quality control. We systematically evaluate multilingual chart understanding with PolyChartQA on state-of-the-art LVLMs and reveal a significant performance gap between English and other languages, particularly low-resource ones. Additionally, we introduce a companion multilingual chart question answering training set, PolyChartQA-Train, on which fine-tuning LVLMs yields substantial gains in multilingual chart understanding across diverse model sizes and architectures. Together, our benchmark provides a foundation for developing globally inclusive vision-language models capable of understanding charts across diverse linguistic contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
Xu, Yichen
Chen, Liangyu
Zhang, Liang
Ma, Jianzhe
Wang, Wenxuan
Jin, Qin
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. To address this limitation, we introduce PolyChartQA, the first large-scale multilingual benchmark for chart question answering, comprising 22,606 charts and 26,151 QA pairs across 10 diverse languages. PolyChartQA is constructed through a scalable pipeline that enables efficient multilingual chart generation via data translation and code reuse, supported by LLM-based translation and rigorous quality control. We systematically evaluate multilingual chart understanding with PolyChartQA on state-of-the-art LVLMs and reveal a significant performance gap between English and other languages, particularly low-resource ones. Additionally, we introduce a companion multilingual chart question answering training set, PolyChartQA-Train, on which fine-tuning LVLMs yields substantial gains in multilingual chart understanding across diverse model sizes and architectures. Together, our benchmark provides a foundation for developing globally inclusive vision-language models capable of understanding charts across diverse linguistic contexts.
title POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2507.11939