An Evaluation-Centric Paradigm for Scientific Visualization Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918143750832128 |
|---|---|
| author | Ai, Kuangshi Miao, Haichao Li, Zhimin Wang, Chaoli Liu, Shusen |
| author_facet | Ai, Kuangshi Miao, Haichao Li, Zhimin Wang, Chaoli Liu, Shusen |
| contents | Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring progress and comparing different agents remains challenging, particularly in scientific visualization (SciVis), due to the absence of comprehensive, large-scale benchmarks for evaluating real-world capabilities. This position paper examines the various types of evaluation required for SciVis agents, outlines the associated challenges, provides a simple proof-of-concept evaluation example, and discusses how evaluation benchmarks can facilitate agent self-improvement. We advocate for a broader collaboration to develop a SciVis agentic evaluation benchmark that would not only assess existing capabilities but also drive innovation and stimulate future development in the field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15160 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Evaluation-Centric Paradigm for Scientific Visualization Agents Ai, Kuangshi Miao, Haichao Li, Zhimin Wang, Chaoli Liu, Shusen Human-Computer Interaction Computation and Language Graphics Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring progress and comparing different agents remains challenging, particularly in scientific visualization (SciVis), due to the absence of comprehensive, large-scale benchmarks for evaluating real-world capabilities. This position paper examines the various types of evaluation required for SciVis agents, outlines the associated challenges, provides a simple proof-of-concept evaluation example, and discusses how evaluation benchmarks can facilitate agent self-improvement. We advocate for a broader collaboration to develop a SciVis agentic evaluation benchmark that would not only assess existing capabilities but also drive innovation and stimulate future development in the field. |
| title | An Evaluation-Centric Paradigm for Scientific Visualization Agents |
| topic | Human-Computer Interaction Computation and Language Graphics |
| url | https://arxiv.org/abs/2509.15160 |