PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jiajun, Zhang, Jianke, Cui, Zeyu, Yang, Jiaxi, Zhang, Lei, Hui, Binyuan, Liu, Qiang, Wang, Zilei, Wang, Liang, Lin, Junyang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918290976145408
author Zhang, Jiajun
Zhang, Jianke
Cui, Zeyu
Yang, Jiaxi
Zhang, Lei
Hui, Binyuan
Liu, Qiang
Wang, Zilei
Wang, Liang
Lin, Junyang
author_facet Zhang, Jiajun
Zhang, Jianke
Cui, Zeyu
Yang, Jiaxi
Zhang, Lei
Hui, Binyuan
Liu, Qiang
Wang, Zilei
Wang, Liang
Lin, Junyang
contents Recent Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation. However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce PlotCraft, a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. The benchmark is structured around seven high-level visualization tasks and encompasses 48 distinct chart types. Crucially, it is the first to systematically evaluate both single-turn generation and multi-turn refinement across a diverse spectrum of task complexities. Our comprehensive evaluation of 23 leading LLMs on PlotCraft reveals obvious performance deficiencies in handling sophisticated visualization tasks. To bridge this performance gap, we develope SynthVis-30K, a large-scale, high-quality dataset of complex visualization code synthesized via a collaborative agent framework. Building upon this dataset, we develope PlotCraftor, a novel code generation model that achieves strong capabilities in complex data visualization with a remarkably small size. Across VisEval, PandasPlotBench, and our proposed PlotCraft, PlotCraftor shows performance comparable to that of leading proprietary approaches. Especially, on hard task, Our model achieves over 50% performance improvement. We will release the benchmark, dataset, and code at https://github.com/Speakn0w/PlotCraft-Benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
Zhang, Jiajun
Zhang, Jianke
Cui, Zeyu
Yang, Jiaxi
Zhang, Lei
Hui, Binyuan
Liu, Qiang
Wang, Zilei
Wang, Liang
Lin, Junyang
Computation and Language
Recent Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation. However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce PlotCraft, a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. The benchmark is structured around seven high-level visualization tasks and encompasses 48 distinct chart types. Crucially, it is the first to systematically evaluate both single-turn generation and multi-turn refinement across a diverse spectrum of task complexities. Our comprehensive evaluation of 23 leading LLMs on PlotCraft reveals obvious performance deficiencies in handling sophisticated visualization tasks. To bridge this performance gap, we develope SynthVis-30K, a large-scale, high-quality dataset of complex visualization code synthesized via a collaborative agent framework. Building upon this dataset, we develope PlotCraftor, a novel code generation model that achieves strong capabilities in complex data visualization with a remarkably small size. Across VisEval, PandasPlotBench, and our proposed PlotCraft, PlotCraftor shows performance comparable to that of leading proprietary approaches. Especially, on hard task, Our model achieves over 50% performance improvement. We will release the benchmark, dataset, and code at https://github.com/Speakn0w/PlotCraft-Benchmark.
title PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
topic Computation and Language
url https://arxiv.org/abs/2511.00010