Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Laskar, Md Tahmid Rahman, Islam, Mohammed Saidul, Mahbub, Ridwan, Masry, Ahmed, Rahman, Mizanur, Bhuiyan, Amran, Nayeem, Mir Tafseer, Joty, Shafiq, Hoque, Enamul, Huang, Jimmy
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916830253154304
author Laskar, Md Tahmid Rahman
Islam, Mohammed Saidul
Mahbub, Ridwan
Masry, Ahmed
Rahman, Mizanur
Bhuiyan, Amran
Nayeem, Mir Tafseer
Joty, Shafiq
Hoque, Enamul
Huang, Jimmy
author_facet Laskar, Md Tahmid Rahman
Islam, Mohammed Saidul
Mahbub, Ridwan
Masry, Ahmed
Rahman, Mizanur
Bhuiyan, Amran
Nayeem, Mir Tafseer
Joty, Shafiq
Hoque, Enamul
Huang, Jimmy
contents Charts are ubiquitous as they help people understand and reason with data. Recently, various downstream tasks, such as chart question answering, chart2text, and fact-checking, have emerged. Large Vision-Language Models (LVLMs) show promise in tackling these tasks, but their evaluation is costly and time-consuming, limiting real-world deployment. While using LVLMs as judges to assess the chart comprehension capabilities of other LVLMs could streamline evaluation processes, challenges like proprietary datasets, restricted access to powerful models, and evaluation costs hinder their adoption in industrial settings. To this end, we present a comprehensive evaluation of 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. We design both pairwise and pointwise evaluation tasks covering criteria like factual correctness, informativeness, and relevancy. Additionally, we analyze LVLM judges based on format adherence, positional consistency, length bias, and instruction-following. We focus on cost-effective LVLMs (<10B parameters) suitable for both research and commercial use, following a standardized evaluation protocol and rubric to measure the LVLM judge's accuracy. Experimental results reveal notable variability: while some open LVLM judges achieve GPT-4-level evaluation performance (about 80% agreement with GPT-4 judgments), others struggle (below ~10% agreement). Our findings highlight that state-of-the-art open-source LVLMs can serve as cost-effective automatic evaluators for chart-related tasks, though biases such as positional preference and length bias persist.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08468
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
Laskar, Md Tahmid Rahman
Islam, Mohammed Saidul
Mahbub, Ridwan
Masry, Ahmed
Rahman, Mizanur
Bhuiyan, Amran
Nayeem, Mir Tafseer
Joty, Shafiq
Hoque, Enamul
Huang, Jimmy
Computation and Language
Computer Vision and Pattern Recognition
Charts are ubiquitous as they help people understand and reason with data. Recently, various downstream tasks, such as chart question answering, chart2text, and fact-checking, have emerged. Large Vision-Language Models (LVLMs) show promise in tackling these tasks, but their evaluation is costly and time-consuming, limiting real-world deployment. While using LVLMs as judges to assess the chart comprehension capabilities of other LVLMs could streamline evaluation processes, challenges like proprietary datasets, restricted access to powerful models, and evaluation costs hinder their adoption in industrial settings. To this end, we present a comprehensive evaluation of 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. We design both pairwise and pointwise evaluation tasks covering criteria like factual correctness, informativeness, and relevancy. Additionally, we analyze LVLM judges based on format adherence, positional consistency, length bias, and instruction-following. We focus on cost-effective LVLMs (<10B parameters) suitable for both research and commercial use, following a standardized evaluation protocol and rubric to measure the LVLM judge's accuracy. Experimental results reveal notable variability: while some open LVLM judges achieve GPT-4-level evaluation performance (about 80% agreement with GPT-4 judgments), others struggle (below ~10% agreement). Our findings highlight that state-of-the-art open-source LVLMs can serve as cost-effective automatic evaluators for chart-related tasks, though biases such as positional preference and length bias persist.
title Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.08468