What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ma, Yan, Zhang, Weiyu, Li, Tianle, Du, Linge, Shen, Xuyang, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
One RL to See Them All: Visual Triple Unified Reinforcement Learning
par: Ma, Yan, et autres
Publié: (2025)
par: Ma, Yan, et autres
Publié: (2025)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
par: Jiang, Dongfu, et autres
Publié: (2025)
par: Jiang, Dongfu, et autres
Publié: (2025)
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
par: Zhang, Yabo, et autres
Publié: (2025)
par: Zhang, Yabo, et autres
Publié: (2025)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
par: Carvalho, Miguel, et autres
Publié: (2025)
par: Carvalho, Miguel, et autres
Publié: (2025)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
par: Yang, Zuhao, et autres
Publié: (2026)
par: Yang, Zuhao, et autres
Publié: (2026)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
par: Shen, Xiaoqian, et autres
Publié: (2025)
par: Shen, Xiaoqian, et autres
Publié: (2025)
OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
par: Fan, Yuxuan, et autres
Publié: (2026)
par: Fan, Yuxuan, et autres
Publié: (2026)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
par: Ma, Yan, et autres
Publié: (2025)
par: Ma, Yan, et autres
Publié: (2025)
Visual Reasoning through Tool-supervised Reinforcement Learning
par: Dong, Qihua, et autres
Publié: (2026)
par: Dong, Qihua, et autres
Publié: (2026)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
par: Huang, Zeyi, et autres
Publié: (2025)
par: Huang, Zeyi, et autres
Publié: (2025)
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
par: Guo, Garvin, et autres
Publié: (2026)
par: Guo, Garvin, et autres
Publié: (2026)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
par: Wang, Chenyu, et autres
Publié: (2024)
par: Wang, Chenyu, et autres
Publié: (2024)
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
par: Yue, Yang, et autres
Publié: (2025)
par: Yue, Yang, et autres
Publié: (2025)
MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
par: Li, Yifan, et autres
Publié: (2025)
par: Li, Yifan, et autres
Publié: (2025)
OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
par: Su, Zhaochen, et autres
Publié: (2025)
par: Su, Zhaochen, et autres
Publié: (2025)
Reinforced Visual Perception with Tools
par: Zhou, Zetong, et autres
Publié: (2025)
par: Zhou, Zetong, et autres
Publié: (2025)
What Really Matters for Learning-based LiDAR-Camera Calibration
par: Huang, Shujuan, et autres
Publié: (2025)
par: Huang, Shujuan, et autres
Publié: (2025)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
par: Zhong, Yaoyao, et autres
Publié: (2023)
par: Zhong, Yaoyao, et autres
Publié: (2023)
AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning
par: Su, Jiaming, et autres
Publié: (2026)
par: Su, Jiaming, et autres
Publié: (2026)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
par: Geigle, Gregor, et autres
Publié: (2024)
par: Geigle, Gregor, et autres
Publié: (2024)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
par: Lu, Meng, et autres
Publié: (2025)
par: Lu, Meng, et autres
Publié: (2025)
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
par: Lyu, Yunbo, et autres
Publié: (2025)
par: Lyu, Yunbo, et autres
Publié: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
par: Ding, Shengyuan, et autres
Publié: (2025)
par: Ding, Shengyuan, et autres
Publié: (2025)
Deep Learning at the Intersection: Certified Robustness as a Tool for 3D Vision
par: S, Gabriel Pérez, et autres
Publié: (2024)
par: S, Gabriel Pérez, et autres
Publié: (2024)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
par: Wang, Yuji, et autres
Publié: (2025)
par: Wang, Yuji, et autres
Publié: (2025)
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
par: Wang, Wenhan, et autres
Publié: (2026)
par: Wang, Wenhan, et autres
Publié: (2026)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
par: Han, Yuhang, et autres
Publié: (2026)
par: Han, Yuhang, et autres
Publié: (2026)
A Calibration Tool for Refractive Underwater Vision
par: Seegräber, Felix, et autres
Publié: (2024)
par: Seegräber, Felix, et autres
Publié: (2024)
Cropper: Vision-Language Model for Image Cropping through In-Context Learning
par: Lee, Seung Hyun, et autres
Publié: (2024)
par: Lee, Seung Hyun, et autres
Publié: (2024)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
par: Jia, Hongrui, et autres
Publié: (2025)
par: Jia, Hongrui, et autres
Publié: (2025)
Learn From Zoom: Decoupled Supervised Contrastive Learning For WCE Image Classification
par: Qiu, Kunpeng, et autres
Publié: (2024)
par: Qiu, Kunpeng, et autres
Publié: (2024)
EventZoom: A Progressive Approach to Event-Based Data Augmentation for Enhanced Neuromorphic Vision
par: Dong, Yiting, et autres
Publié: (2024)
par: Dong, Yiting, et autres
Publié: (2024)
Does the Skeleton-Recall Loss Really Work?
par: Arora, Devansh, et autres
Publié: (2025)
par: Arora, Devansh, et autres
Publié: (2025)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
par: Kumar, Sunil, et autres
Publié: (2025)
par: Kumar, Sunil, et autres
Publié: (2025)
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
par: Shen, Haozhan, et autres
Publié: (2024)
par: Shen, Haozhan, et autres
Publié: (2024)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
par: Jeddi, Ahmadreza, et autres
Publié: (2026)
par: Jeddi, Ahmadreza, et autres
Publié: (2026)
Reliable Disentanglement Multi-view Learning Against View Adversarial Attacks
par: Wang, Xuyang, et autres
Publié: (2025)
par: Wang, Xuyang, et autres
Publié: (2025)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
par: Wang, Chaoyang, et autres
Publié: (2025)
par: Wang, Chaoyang, et autres
Publié: (2025)
On the Global Photometric Alignment for Low-Level Vision
par: Li, Mingjia, et autres
Publié: (2026)
par: Li, Mingjia, et autres
Publié: (2026)
Documents similaires
-
One RL to See Them All: Visual Triple Unified Reinforcement Learning
par: Ma, Yan, et autres
Publié: (2025) -
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
par: Jiang, Dongfu, et autres
Publié: (2025) -
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
par: Zhang, Yabo, et autres
Publié: (2025) -
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
par: Carvalho, Miguel, et autres
Publié: (2025) -
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
par: Yang, Zuhao, et autres
Publié: (2026)