The Plot Thickens: Quantitative Part-by-Part Exploration of MLLM Visualization Literacy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Valentim, Matheus, Dhanoa, Vaishali, León, Gabriela Molina, Elmqvist, Niklas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916672604995584
author Valentim, Matheus
Dhanoa, Vaishali
León, Gabriela Molina
Elmqvist, Niklas
author_facet Valentim, Matheus
Dhanoa, Vaishali
León, Gabriela Molina
Elmqvist, Niklas
contents Multimodal Large Language Models (MLLMs) can interpret data visualizations, but what makes a visualization understandable to these models? Do factors like color, shape, and text influence legibility, and how does this compare to human perception? In this paper, we build on prior work to systematically assess which visualization characteristics impact MLLM interpretability. We expanded the Visualization Literacy Assessment Test (VLAT) test set from 12 to 380 visualizations by varying plot types, colors, and titles. This allowed us to statistically analyze how these features affect model performance. Our findings suggest that while color palettes have no significant impact on accuracy, plot types and the type of title significantly affect MLLM performance. We observe similar trends for model omissions. Based on these insights, we look into which plot types are beneficial for MLLMs in different tasks and propose visualization design principles that enhance MLLM readability. Additionally, we make the extended VLAT test set, VLAT ex, publicly available on https://osf.io/ermwx/ together with our supplemental material for future model testing and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Plot Thickens: Quantitative Part-by-Part Exploration of MLLM Visualization Literacy
Valentim, Matheus
Dhanoa, Vaishali
León, Gabriela Molina
Elmqvist, Niklas
Human-Computer Interaction
Multimodal Large Language Models (MLLMs) can interpret data visualizations, but what makes a visualization understandable to these models? Do factors like color, shape, and text influence legibility, and how does this compare to human perception? In this paper, we build on prior work to systematically assess which visualization characteristics impact MLLM interpretability. We expanded the Visualization Literacy Assessment Test (VLAT) test set from 12 to 380 visualizations by varying plot types, colors, and titles. This allowed us to statistically analyze how these features affect model performance. Our findings suggest that while color palettes have no significant impact on accuracy, plot types and the type of title significantly affect MLLM performance. We observe similar trends for model omissions. Based on these insights, we look into which plot types are beneficial for MLLMs in different tasks and propose visualization design principles that enhance MLLM readability. Additionally, we make the extended VLAT test set, VLAT ex, publicly available on https://osf.io/ermwx/ together with our supplemental material for future model testing and evaluation.
title The Plot Thickens: Quantitative Part-by-Part Exploration of MLLM Visualization Literacy
topic Human-Computer Interaction
url https://arxiv.org/abs/2504.02217