Exploring Interaction Paradigms for LLM Agents in Scientific Visualization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vonderhorst, Jackson, Ai, Kuangshi, Miao, Haichao, Liu, Shusen, Wang, Chaoli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909038907752448
author Vonderhorst, Jackson
Ai, Kuangshi
Miao, Haichao
Liu, Shusen
Wang, Chaoli
author_facet Vonderhorst, Jackson
Ai, Kuangshi
Miao, Haichao
Liu, Shusen
Wang, Chaoli
contents This paper examines how different types of large language model (LLM) agents perform on scientific visualization (SciVis) tasks, where users generate visualization workflows from natural-language instructions. We compare three primary interaction paradigms, including domain-specific agents with structured tool use, computer-use agents, and general-purpose coding agents, by evaluating eight representative agents across 15 benchmark tasks and measuring visualization quality, efficiency, robustness, and computational cost. We further analyze interaction modalities, including code scripts and model context protocol (MCP) or API calls for structured tool use, as well as command-line interfaces (CLI) and graphical user interfaces (GUI) for more general interaction, while additionally studying the effect of persistent memory in selected agents. The results reveal clear tradeoffs across paradigms and modalities. General-purpose coding agents achieve the highest task success rates but are computationally expensive, while domain-specific agents are more efficient and stable but less flexible. Computer-use agents perform well on individual steps but struggle with longer multi-step workflows, indicating that long-horizon planning is their primary limitation. Across both CLI- and GUI-based settings, persistent memory improves performance over repeated trials, although its benefits depend on the underlying interaction mode and the quality of feedback. These findings suggest that no single approach is sufficient, and future SciVis systems should combine structured tool use, interactive capabilities, and adaptive memory mechanisms to balance performance, robustness, and flexibility.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27996
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploring Interaction Paradigms for LLM Agents in Scientific Visualization
Vonderhorst, Jackson
Ai, Kuangshi
Miao, Haichao
Liu, Shusen
Wang, Chaoli
Artificial Intelligence
Graphics
Human-Computer Interaction
This paper examines how different types of large language model (LLM) agents perform on scientific visualization (SciVis) tasks, where users generate visualization workflows from natural-language instructions. We compare three primary interaction paradigms, including domain-specific agents with structured tool use, computer-use agents, and general-purpose coding agents, by evaluating eight representative agents across 15 benchmark tasks and measuring visualization quality, efficiency, robustness, and computational cost. We further analyze interaction modalities, including code scripts and model context protocol (MCP) or API calls for structured tool use, as well as command-line interfaces (CLI) and graphical user interfaces (GUI) for more general interaction, while additionally studying the effect of persistent memory in selected agents. The results reveal clear tradeoffs across paradigms and modalities. General-purpose coding agents achieve the highest task success rates but are computationally expensive, while domain-specific agents are more efficient and stable but less flexible. Computer-use agents perform well on individual steps but struggle with longer multi-step workflows, indicating that long-horizon planning is their primary limitation. Across both CLI- and GUI-based settings, persistent memory improves performance over repeated trials, although its benefits depend on the underlying interaction mode and the quality of feedback. These findings suggest that no single approach is sufficient, and future SciVis systems should combine structured tool use, interactive capabilities, and adaptive memory mechanisms to balance performance, robustness, and flexibility.
title Exploring Interaction Paradigms for LLM Agents in Scientific Visualization
topic Artificial Intelligence
Graphics
Human-Computer Interaction
url https://arxiv.org/abs/2604.27996