VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Zeyi, Ji, Yuyang, Rajan, Anirudh Sundara, Cai, Zefan, Xiao, Wen, Wang, Haohan, Hu, Junjie, Lee, Yong Jae
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912492238667776
author Huang, Zeyi
Ji, Yuyang
Rajan, Anirudh Sundara
Cai, Zefan
Xiao, Wen
Wang, Haohan
Hu, Junjie
Lee, Yong Jae
author_facet Huang, Zeyi
Ji, Yuyang
Rajan, Anirudh Sundara
Cai, Zefan
Xiao, Wen
Wang, Haohan
Hu, Junjie
Lee, Yong Jae
contents We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning either rely on training-free prompting or large-scale fine-tuning; both lack active tool exploration and typically assume limited tool diversity, and fine-tuning methods additionally demand extensive human supervision. In contrast, VisTA leverages end-to-end reinforcement learning to iteratively refine sophisticated, query-specific tool selection strategies, using task outcomes as feedback signals. Through Group Relative Policy Optimization (GRPO), our framework enables an agent to autonomously discover effective tool-selection pathways without requiring explicit reasoning supervision. Experiments on the ChartQA, Geometry3K, and BlindTest benchmarks demonstrate that VisTA achieves substantial performance gains over training-free baselines, especially on out-of-distribution examples. These results highlight VisTA's ability to enhance generalization, adaptively utilize diverse tools, and pave the way for flexible, experience-driven visual reasoning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20289
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Huang, Zeyi
Ji, Yuyang
Rajan, Anirudh Sundara
Cai, Zefan
Xiao, Wen
Wang, Haohan
Hu, Junjie
Lee, Yong Jae
Computer Vision and Pattern Recognition
We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning either rely on training-free prompting or large-scale fine-tuning; both lack active tool exploration and typically assume limited tool diversity, and fine-tuning methods additionally demand extensive human supervision. In contrast, VisTA leverages end-to-end reinforcement learning to iteratively refine sophisticated, query-specific tool selection strategies, using task outcomes as feedback signals. Through Group Relative Policy Optimization (GRPO), our framework enables an agent to autonomously discover effective tool-selection pathways without requiring explicit reasoning supervision. Experiments on the ChartQA, Geometry3K, and BlindTest benchmarks demonstrate that VisTA achieves substantial performance gains over training-free baselines, especially on out-of-distribution examples. These results highlight VisTA's ability to enhance generalization, adaptively utilize diverse tools, and pave the way for flexible, experience-driven visual reasoning systems.
title VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20289