CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bharti, Shubham, Cheng, Shiyun, Rho, Jihyun, Zhang, Jianrui, Cai, Mu, Lee, Yong Jae, Rau, Martina, Zhu, Xiaojin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916814024343552
author Bharti, Shubham
Cheng, Shiyun
Rho, Jihyun
Zhang, Jianrui
Cai, Mu
Lee, Yong Jae
Rau, Martina
Zhu, Xiaojin
author_facet Bharti, Shubham
Cheng, Shiyun
Rho, Jihyun
Zhang, Jianrui
Cai, Mu
Lee, Yong Jae
Rau, Martina
Zhu, Xiaojin
contents We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizations though charts. CHARTOM consists of carefully designed charts and associated questions that require a language model to not only correctly comprehend the factual content in the chart (the FACT question) but also judge whether the chart will be misleading to a human readers (the MIND question), a dual capability with significant societal benefits. We detail the construction of our benchmark including its calibration on human performance and estimation of MIND ground truth called the Human Misleadingness Index. We evaluated several leading LLMs -- including GPT, Claude, Gemini, Qwen, Llama, and Llava series models -- on the CHARTOM dataset and found that it was challenging to all models both on FACT and MIND questions. This highlights the limitations of current LLMs and presents significant opportunity for future LLMs to improve on understanding misleading charts.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14419
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
Bharti, Shubham
Cheng, Shiyun
Rho, Jihyun
Zhang, Jianrui
Cai, Mu
Lee, Yong Jae
Rau, Martina
Zhu, Xiaojin
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizations though charts. CHARTOM consists of carefully designed charts and associated questions that require a language model to not only correctly comprehend the factual content in the chart (the FACT question) but also judge whether the chart will be misleading to a human readers (the MIND question), a dual capability with significant societal benefits. We detail the construction of our benchmark including its calibration on human performance and estimation of MIND ground truth called the Human Misleadingness Index. We evaluated several leading LLMs -- including GPT, Claude, Gemini, Qwen, Llama, and Llava series models -- on the CHARTOM dataset and found that it was challenging to all models both on FACT and MIND questions. This highlights the limitations of current LLMs and presents significant opportunity for future LLMs to improve on understanding misleading charts.
title CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.14419