Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seon Gyeom, Choi, Jae Young, Rossi, Ryan, Koh, Eunyee, Lee, Tak Yeon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908376194088960
author Kim, Seon Gyeom
Choi, Jae Young
Rossi, Ryan
Koh, Eunyee
Lee, Tak Yeon
author_facet Kim, Seon Gyeom
Choi, Jae Young
Rossi, Ryan
Koh, Eunyee
Lee, Tak Yeon
contents The field of Multimodal Large Language Models (MLLMs) has made remarkable progress in visual understanding tasks, presenting a vast opportunity to predict the perceptual and emotional impact of charts. However, it also raises concerns, as many applications of LLMs are based on overgeneralized assumptions from a few examples, lacking sufficient validation of their performance and effectiveness. We introduce Chart-to-Experience, a benchmark dataset comprising 36 charts, evaluated by crowdsourced workers for their impact on seven experiential factors. Using the dataset as ground truth, we evaluated capabilities of state-of-the-art MLLMs on two tasks: direct prediction and pairwise comparison of charts. Our findings imply that MLLMs are not as sensitive as human evaluators when assessing individual charts, but are accurate and reliable in pairwise comparisons.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
Kim, Seon Gyeom
Choi, Jae Young
Rossi, Ryan
Koh, Eunyee
Lee, Tak Yeon
Human-Computer Interaction
Computation and Language
The field of Multimodal Large Language Models (MLLMs) has made remarkable progress in visual understanding tasks, presenting a vast opportunity to predict the perceptual and emotional impact of charts. However, it also raises concerns, as many applications of LLMs are based on overgeneralized assumptions from a few examples, lacking sufficient validation of their performance and effectiveness. We introduce Chart-to-Experience, a benchmark dataset comprising 36 charts, evaluated by crowdsourced workers for their impact on seven experiential factors. Using the dataset as ground truth, we evaluated capabilities of state-of-the-art MLLMs on two tasks: direct prediction and pairwise comparison of charts. Our findings imply that MLLMs are not as sensitive as human evaluators when assessing individual charts, but are accurate and reliable in pairwise comparisons.
title Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2505.17374