CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Verma, Arnav, Mukherjee, Kushin, Potts, Christopher, Kreiss, Elisa, Fan, Judith E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913854787682304
author Verma, Arnav
Mukherjee, Kushin
Potts, Christopher
Kreiss, Elisa
Fan, Judith E.
author_facet Verma, Arnav
Mukherjee, Kushin
Potts, Christopher
Kreiss, Elisa
Fan, Judith E.
contents Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs arranged in a conventionalized format one has previously learned to parse. Recently developed vision-language models are, in principle, promising candidates for developing computational models of these cognitive operations. However, it is currently unclear to what degree these models emulate human behavior on tasks that involve reasoning about data visualizations. This gap reflects limitations in prior work that has evaluated data visualization understanding in artificial systems using measures that differ from those typically used to assess these abilities in humans. Here we evaluated eight vision-language models on six data visualization literacy assessments designed for humans and compared model responses to those of human participants. We found that these models performed worse than human participants on average, and this performance gap persisted even when using relatively lenient criteria to assess model performance. Moreover, while relative performance across items was somewhat correlated between models and humans, all models produced patterns of errors that were reliably distinct from those produced by human participants. Taken together, these findings suggest significant opportunities for further development of artificial systems that might serve as useful models of how humans reason about data visualizations. All code and data needed to reproduce these results are available at: https://osf.io/e25mu/?view_only=399daff5a14d4b16b09473cf19043f18.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
Verma, Arnav
Mukherjee, Kushin
Potts, Christopher
Kreiss, Elisa
Fan, Judith E.
Human-Computer Interaction
Computation and Language
Computer Vision and Pattern Recognition
Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs arranged in a conventionalized format one has previously learned to parse. Recently developed vision-language models are, in principle, promising candidates for developing computational models of these cognitive operations. However, it is currently unclear to what degree these models emulate human behavior on tasks that involve reasoning about data visualizations. This gap reflects limitations in prior work that has evaluated data visualization understanding in artificial systems using measures that differ from those typically used to assess these abilities in humans. Here we evaluated eight vision-language models on six data visualization literacy assessments designed for humans and compared model responses to those of human participants. We found that these models performed worse than human participants on average, and this performance gap persisted even when using relatively lenient criteria to assess model performance. Moreover, while relative performance across items was somewhat correlated between models and humans, all models produced patterns of errors that were reliably distinct from those produced by human participants. Taken together, these findings suggest significant opportunities for further development of artificial systems that might serve as useful models of how humans reason about data visualizations. All code and data needed to reproduce these results are available at: https://osf.io/e25mu/?view_only=399daff5a14d4b16b09473cf19043f18.
title CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
topic Human-Computer Interaction
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17202