VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zeyi, Ji, Yuyang, Rajan, Anirudh Sundara, Cai, Zefan, Xiao, Wen, Wang, Haohan, Hu, Junjie, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
by: Rajan, Anirudh Sundara, et al.
Published: (2025)
by: Rajan, Anirudh Sundara, et al.
Published: (2025)
VisTA-SR: Improving the Accuracy and Resolution of Low-Cost Thermal Imaging Cameras for Agriculture
by: Yun, Heesup, et al.
Published: (2024)
by: Yun, Heesup, et al.
Published: (2024)
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
by: Rajan, Anirudh Sundara, et al.
Published: (2026)
by: Rajan, Anirudh Sundara, et al.
Published: (2026)
Aligned Datasets Improve Detection of Latent Diffusion-Generated Images
by: Rajan, Anirudh Sundara, et al.
Published: (2024)
by: Rajan, Anirudh Sundara, et al.
Published: (2024)
Reinforced Visual Perception with Tools
by: Zhou, Zetong, et al.
Published: (2025)
by: Zhou, Zetong, et al.
Published: (2025)
VisTA: Vision-Text Alignment Model with Contrastive Learning using Multimodal Data for Evidence-Driven, Reliable, and Explainable Alzheimer's Disease Diagnosis
by: Can, Duy-Cat, et al.
Published: (2025)
by: Can, Duy-Cat, et al.
Published: (2025)
Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding
by: Ji, Yuyang, et al.
Published: (2025)
by: Ji, Yuyang, et al.
Published: (2025)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Understanding Expert Exploration in EHR Visualization Tools: The ParcoursVis Use Case
by: Assor, Ambre, et al.
Published: (2025)
by: Assor, Ambre, et al.
Published: (2025)
VisAgent: Narrative-Preserving Story Visualization Framework
by: Kim, Seungkwon, et al.
Published: (2025)
by: Kim, Seungkwon, et al.
Published: (2025)
Visual Reasoning through Tool-supervised Reinforcement Learning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Experts
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents
by: Ta, Anh, et al.
Published: (2026)
by: Ta, Anh, et al.
Published: (2026)
Delta Attention Residuals
by: Luo, Cheng, et al.
Published: (2026)
by: Luo, Cheng, et al.
Published: (2026)
VESTA: Visual Exploration with Statistical Tool Agents
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
by: Mattioli, Gabriele, et al.
Published: (2026)
by: Mattioli, Gabriele, et al.
Published: (2026)
ReVis: Towards Reusable Image-Based Visualizations with MLLMs
by: Wen, Xiaolin, et al.
Published: (2026)
by: Wen, Xiaolin, et al.
Published: (2026)
The Visual Debugger Tool
by: Kräuter, Tim, et al.
Published: (2024)
by: Kräuter, Tim, et al.
Published: (2024)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
by: Khaki, Saeed, et al.
Published: (2026)
by: Khaki, Saeed, et al.
Published: (2026)
VisCoder2: Building Multi-Language Visualization Coding Agents
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
A Practical Tool for Visualizing and Measuring Model Selection Uncertainty
by: Sheng Ren, et al.
Published: (2025)
by: Sheng Ren, et al.
Published: (2025)
Visual-RFT: Visual Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
COMMA: A Communicative Multimodal Multi-Agent Benchmark
by: Ossowski, Timothy, et al.
Published: (2024)
by: Ossowski, Timothy, et al.
Published: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
by: Zhang, Situo, et al.
Published: (2026)
by: Zhang, Situo, et al.
Published: (2026)
InTraVisTo: Inside Transformer Visualisation Tool
by: Brunello, Nicolò, et al.
Published: (2025)
by: Brunello, Nicolò, et al.
Published: (2025)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
A Visualized Framework for Event Cooperation with Generative Agents
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
by: Ding, Keyan, et al.
Published: (2025)
by: Ding, Keyan, et al.
Published: (2025)
ParaView-MCP: An Autonomous Visualization Agent with Direct Tool Use
by: Liu, Shusen, et al.
Published: (2025)
by: Liu, Shusen, et al.
Published: (2025)
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
by: Liao, Kunpeng, et al.
Published: (2026)
by: Liao, Kunpeng, et al.
Published: (2026)
RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models
by: Wang, Zefan, et al.
Published: (2023)
by: Wang, Zefan, et al.
Published: (2023)
ChatVis: Large Language Model Agent for Generating Scientific Visualizations
by: Peterka, Tom, et al.
Published: (2025)
by: Peterka, Tom, et al.
Published: (2025)
Visualization for Human-Centered AI Tools
by: Hoque, Md Naimul, et al.
Published: (2024)
by: Hoque, Md Naimul, et al.
Published: (2024)
Similar Items
-
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
by: Rajan, Anirudh Sundara, et al.
Published: (2025) -
VisTA-SR: Improving the Accuracy and Resolution of Low-Cost Thermal Imaging Cameras for Agriculture
by: Yun, Heesup, et al.
Published: (2024) -
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
by: Rajan, Anirudh Sundara, et al.
Published: (2026) -
Aligned Datasets Improve Detection of Latent Diffusion-Generated Images
by: Rajan, Anirudh Sundara, et al.
Published: (2024) -
Reinforced Visual Perception with Tools
by: Zhou, Zetong, et al.
Published: (2025)