VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Dongfu, Lu, Yi, Li, Zhuofeng, Lyu, Zhiheng, Nie, Ping, Wang, Haozhe, Su, Alex, Chen, Hui, Zou, Kai, Du, Chao, Pang, Tianyu, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
von: Ruan, Chi, et al.
Veröffentlicht: (2026)
von: Ruan, Chi, et al.
Veröffentlicht: (2026)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
von: Zeng, Huaye, et al.
Veröffentlicht: (2025)
von: Zeng, Huaye, et al.
Veröffentlicht: (2025)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
von: Li, Zhuofeng, et al.
Veröffentlicht: (2026)
von: Li, Zhuofeng, et al.
Veröffentlicht: (2026)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
von: Ruan, Chi, et al.
Veröffentlicht: (2025)
von: Ruan, Chi, et al.
Veröffentlicht: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
von: Li, Zhuofeng, et al.
Veröffentlicht: (2025)
von: Li, Zhuofeng, et al.
Veröffentlicht: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
ToolRM: Towards Agentic Tool-Use Reward Modeling
von: Li, Renhao, et al.
Veröffentlicht: (2025)
von: Li, Renhao, et al.
Veröffentlicht: (2025)
Agentic Tool Use in Large Language Models
von: Hu, Jinchao, et al.
Veröffentlicht: (2026)
von: Hu, Jinchao, et al.
Veröffentlicht: (2026)
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
von: Lu, Pan, et al.
Veröffentlicht: (2025)
von: Lu, Pan, et al.
Veröffentlicht: (2025)
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
von: Zhang, Yabo, et al.
Veröffentlicht: (2025)
von: Zhang, Yabo, et al.
Veröffentlicht: (2025)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
RewardHarness: Self-Evolving Agentic Post-Training
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents
von: Li, Zhuofeng, et al.
Veröffentlicht: (2026)
von: Li, Zhuofeng, et al.
Veröffentlicht: (2026)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
von: Yang, Zuhao, et al.
Veröffentlicht: (2026)
von: Yang, Zuhao, et al.
Veröffentlicht: (2026)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
von: Ku, Max, et al.
Veröffentlicht: (2023)
von: Ku, Max, et al.
Veröffentlicht: (2023)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
von: Lyu, Yuanjie, et al.
Veröffentlicht: (2026)
von: Lyu, Yuanjie, et al.
Veröffentlicht: (2026)
VisCoder2: Building Multi-Language Visualization Coding Agents
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
von: Feng, Jiazhan, et al.
Veröffentlicht: (2025)
von: Feng, Jiazhan, et al.
Veröffentlicht: (2025)
GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
von: Lu, Meng, et al.
Veröffentlicht: (2025)
von: Lu, Meng, et al.
Veröffentlicht: (2025)
General-Reasoner: Advancing LLM Reasoning Across All Domains
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
PyVision: Agentic Vision with Dynamic Tooling
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
von: Yan, Shilin, et al.
Veröffentlicht: (2026)
von: Yan, Shilin, et al.
Veröffentlicht: (2026)
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
von: Huang, Shouzheng, et al.
Veröffentlicht: (2026)
von: Huang, Shouzheng, et al.
Veröffentlicht: (2026)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
von: Lei, Fei, et al.
Veröffentlicht: (2025)
von: Lei, Fei, et al.
Veröffentlicht: (2025)
BAFFLE: A Baseline of Backpropagation-Free Federated Learning
von: Feng, Haozhe, et al.
Veröffentlicht: (2023)
von: Feng, Haozhe, et al.
Veröffentlicht: (2023)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
Reinforced Visual Perception with Tools
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025) -
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
von: Ruan, Chi, et al.
Veröffentlicht: (2026) -
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
von: Zeng, Huaye, et al.
Veröffentlicht: (2025) -
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
von: Li, Zhuofeng, et al.
Veröffentlicht: (2026) -
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
von: Ruan, Chi, et al.
Veröffentlicht: (2025)