Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Meng, Xu, Ran, Fang, Yi, Zhang, Wenxuan, Yu, Yue, Srivastava, Gaurav, Zhuang, Yuchen, Elhoseiny, Mohamed, Fleming, Charles, Yang, Carl, Tu, Zhengzhong, Xie, Yang, Xiao, Guanghua, Wang, Hanrui, Jin, Di, Shi, Wenqi, Wang, Xuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908673396178944
author Lu, Meng
Xu, Ran
Fang, Yi
Zhang, Wenxuan
Yu, Yue
Srivastava, Gaurav
Zhuang, Yuchen
Elhoseiny, Mohamed
Fleming, Charles
Yang, Carl
Tu, Zhengzhong
Xie, Yang
Xiao, Guanghua
Wang, Hanrui
Jin, Di
Shi, Wenqi
Wang, Xuan
author_facet Lu, Meng
Xu, Ran
Fang, Yi
Zhang, Wenxuan
Yu, Yue
Srivastava, Gaurav
Zhuang, Yuchen
Elhoseiny, Mohamed
Fleming, Charles
Yang, Carl
Tu, Zhengzhong
Xie, Yang
Xiao, Guanghua
Wang, Hanrui
Jin, Di
Shi, Wenqi
Wang, Xuan
contents While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images", i.e., to reason through multi-step visual interactions, remains limited. We introduce VISTA-Gym, a scalable training environment for incentivizing tool-integrated visual reasoning capabilities in VLMs. VISTA-Gym unifies diverse real-world multimodal reasoning tasks (7 tasks from 13 datasets in total) with a standardized interface for visual tools (e.g., grounding, parsing), executable interaction loops, verifiable feedback signals, and efficient trajectory logging, enabling visual agentic reinforcement learning at scale. While recent VLMs exhibit strong text-only reasoning, both proprietary and open-source models still struggle with tool selection, invocation, and coordination. With VISTA-Gym, we train VISTA-R1 to interleave tool-use with agentic reasoning via multi-turn trajectory sampling and end-to-end reinforcement learning. Extensive experiments across 11 public reasoning-intensive VQA benchmarks show that VISTA-R1-8B outperforms state-of-the-art baselines with similar sizes by 9.51%-18.72%, demonstrating VISTA-Gym as an effective training ground to unlock the tool-integrated reasoning capabilities for VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
Lu, Meng
Xu, Ran
Fang, Yi
Zhang, Wenxuan
Yu, Yue
Srivastava, Gaurav
Zhuang, Yuchen
Elhoseiny, Mohamed
Fleming, Charles
Yang, Carl
Tu, Zhengzhong
Xie, Yang
Xiao, Guanghua
Wang, Hanrui
Jin, Di
Shi, Wenqi
Wang, Xuan
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images", i.e., to reason through multi-step visual interactions, remains limited. We introduce VISTA-Gym, a scalable training environment for incentivizing tool-integrated visual reasoning capabilities in VLMs. VISTA-Gym unifies diverse real-world multimodal reasoning tasks (7 tasks from 13 datasets in total) with a standardized interface for visual tools (e.g., grounding, parsing), executable interaction loops, verifiable feedback signals, and efficient trajectory logging, enabling visual agentic reinforcement learning at scale. While recent VLMs exhibit strong text-only reasoning, both proprietary and open-source models still struggle with tool selection, invocation, and coordination. With VISTA-Gym, we train VISTA-R1 to interleave tool-use with agentic reasoning via multi-turn trajectory sampling and end-to-end reinforcement learning. Extensive experiments across 11 public reasoning-intensive VQA benchmarks show that VISTA-R1-8B outperforms state-of-the-art baselines with similar sizes by 9.51%-18.72%, demonstrating VISTA-Gym as an effective training ground to unlock the tool-integrated reasoning capabilities for VLMs.
title Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19773