Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Zhen, Weng, Luoxuan, Tang, Yinghao, Zhang, Runjin, Liu, Yuxin, Pan, Bo, Zhu, Minfeng, Chen, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917990041124864
author Wen, Zhen
Weng, Luoxuan
Tang, Yinghao
Zhang, Runjin
Liu, Yuxin
Pan, Bo
Zhu, Minfeng
Chen, Wei
author_facet Wen, Zhen
Weng, Luoxuan
Tang, Yinghao
Zhang, Runjin
Liu, Yuxin
Pan, Bo
Zhu, Minfeng
Chen, Wei
contents Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. Through two case studies and a controlled user study, we demonstrate that VisPilot provides a more intuitive way to create visualizations without affecting the overall task efficiency compared to text-only prompting approaches. Furthermore, we analyze the impact of text and visual prompts in different visualization tasks. Our findings highlight the importance of multimodal prompting in improving the usability of LLMs for visualization authoring. We discuss design implications for future visualization systems and provide insights into how multimodal prompts can enhance human-AI collaboration in creative visualization tasks. All materials are available at https://OSF.IO/2QRAK.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Multimodal Prompt for Visualization Authoring with Large Language Models
Wen, Zhen
Weng, Luoxuan
Tang, Yinghao
Zhang, Runjin
Liu, Yuxin
Pan, Bo
Zhu, Minfeng
Chen, Wei
Human-Computer Interaction
Artificial Intelligence
Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. Through two case studies and a controlled user study, we demonstrate that VisPilot provides a more intuitive way to create visualizations without affecting the overall task efficiency compared to text-only prompting approaches. Furthermore, we analyze the impact of text and visual prompts in different visualization tasks. Our findings highlight the importance of multimodal prompting in improving the usability of LLMs for visualization authoring. We discuss design implications for future visualization systems and provide insights into how multimodal prompts can enhance human-AI collaboration in creative visualization tasks. All materials are available at https://OSF.IO/2QRAK.
title Exploring Multimodal Prompt for Visualization Authoring with Large Language Models
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2504.13700