An Evaluation-Centric Paradigm for Scientific Visualization Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ai, Kuangshi, Miao, Haichao, Li, Zhimin, Wang, Chaoli, Liu, Shusen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918143750832128
author Ai, Kuangshi
Miao, Haichao
Li, Zhimin
Wang, Chaoli
Liu, Shusen
author_facet Ai, Kuangshi
Miao, Haichao
Li, Zhimin
Wang, Chaoli
Liu, Shusen
contents Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring progress and comparing different agents remains challenging, particularly in scientific visualization (SciVis), due to the absence of comprehensive, large-scale benchmarks for evaluating real-world capabilities. This position paper examines the various types of evaluation required for SciVis agents, outlines the associated challenges, provides a simple proof-of-concept evaluation example, and discusses how evaluation benchmarks can facilitate agent self-improvement. We advocate for a broader collaboration to develop a SciVis agentic evaluation benchmark that would not only assess existing capabilities but also drive innovation and stimulate future development in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Evaluation-Centric Paradigm for Scientific Visualization Agents
Ai, Kuangshi
Miao, Haichao
Li, Zhimin
Wang, Chaoli
Liu, Shusen
Human-Computer Interaction
Computation and Language
Graphics
Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring progress and comparing different agents remains challenging, particularly in scientific visualization (SciVis), due to the absence of comprehensive, large-scale benchmarks for evaluating real-world capabilities. This position paper examines the various types of evaluation required for SciVis agents, outlines the associated challenges, provides a simple proof-of-concept evaluation example, and discusses how evaluation benchmarks can facilitate agent self-improvement. We advocate for a broader collaboration to develop a SciVis agentic evaluation benchmark that would not only assess existing capabilities but also drive innovation and stimulate future development in the field.
title An Evaluation-Centric Paradigm for Scientific Visualization Agents
topic Human-Computer Interaction
Computation and Language
Graphics
url https://arxiv.org/abs/2509.15160