SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Zhenghai, Zheng, Longtao, Liu, Qian, Li, Yingru, Zheng, Xiaosen, Ma, Zejun, An, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026)
by: Li, Yingru, et al.
Published: (2026)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
Simple, unified analysis of Johnson-Lindenstrauss with applications
by: Li, Yingru
Published: (2024)
by: Li, Yingru
Published: (2024)
MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications
by: Vinayagame, Vishnou, et al.
Published: (2024)
by: Vinayagame, Vishnou, et al.
Published: (2024)
S$^2$AC: Energy-Based Reinforcement Learning with Stein Soft Actor Critic
by: Messaoud, Safa, et al.
Published: (2024)
by: Messaoud, Safa, et al.
Published: (2024)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
by: Zeng, Weihao, et al.
Published: (2025)
by: Zeng, Weihao, et al.
Published: (2025)
Group-in-Group Policy Optimization for LLM Agent Training
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
Probability Tools for Sequential Random Projection
by: Li, Yingru
Published: (2024)
by: Li, Yingru
Published: (2024)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
by: Liu, Licheng, et al.
Published: (2025)
by: Liu, Licheng, et al.
Published: (2025)
AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender Systems
by: Xue, Zhenghai, et al.
Published: (2023)
by: Xue, Zhenghai, et al.
Published: (2023)
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning
by: Xue, Zhenghai, et al.
Published: (2025)
by: Xue, Zhenghai, et al.
Published: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
by: Lu, Yuhang, et al.
Published: (2025)
by: Lu, Yuhang, et al.
Published: (2025)
An End-to-End Deep Reinforcement Learning Approach for Solving the Traveling Salesman Problem with Drones
by: Zeng, Taihelong, et al.
Published: (2025)
by: Zeng, Taihelong, et al.
Published: (2025)
End-to-End Anti-Backdoor Learning on Images and Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
by: Chen, Guanzhong, et al.
Published: (2025)
by: Chen, Guanzhong, et al.
Published: (2025)
State Regularized Policy Optimization on Data with Dynamics Shift
by: Xue, Zhenghai, et al.
Published: (2023)
by: Xue, Zhenghai, et al.
Published: (2023)
End-to-End Framework Integrating Generative AI and Deep Reinforcement Learning for Autonomous Ultrasound Scanning
by: Elmekki, Hanae, et al.
Published: (2025)
by: Elmekki, Hanae, et al.
Published: (2025)
End-to-End Modeling Hierarchical Time Series Using Autoregressive Transformer and Conditional Normalizing Flow based Reconciliation
by: Wang, Shiyu, et al.
Published: (2022)
by: Wang, Shiyu, et al.
Published: (2022)
Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
End-to-End Reinforcement Learning for Torque Based Variable Height Hopping
by: Soni, Raghav, et al.
Published: (2023)
by: Soni, Raghav, et al.
Published: (2023)
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
by: Zhang, Kehao, et al.
Published: (2026)
by: Zhang, Kehao, et al.
Published: (2026)
An End-to-End Reinforcement Learning Based Approach for Micro-View Order-Dispatching in Ride-Hailing
by: Yue, Xinlang, et al.
Published: (2024)
by: Yue, Xinlang, et al.
Published: (2024)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
by: Zhang, Yaxiang, et al.
Published: (2026)
by: Zhang, Yaxiang, et al.
Published: (2026)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
by: Hanneke, Steve, et al.
Published: (2026)
by: Hanneke, Steve, et al.
Published: (2026)
End-to-End Reinforcement Learning of Curative Curtailment with Partial Measurement Availability
by: Wolf, Hinrikus, et al.
Published: (2024)
by: Wolf, Hinrikus, et al.
Published: (2024)
MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair
by: Li, Changqing, et al.
Published: (2025)
by: Li, Changqing, et al.
Published: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
End-to-End Reinforcement Learning of Koopman Models for eNMPC of an Air Separation Unit
by: Mayfrank, Daniel, et al.
Published: (2025)
by: Mayfrank, Daniel, et al.
Published: (2025)
End-to-End Reinforcement Learning of Koopman Models for Economic Nonlinear Model Predictive Control
by: Mayfrank, Daniel, et al.
Published: (2023)
by: Mayfrank, Daniel, et al.
Published: (2023)
Semi-Supervised End-To-End Contrastive Learning For Time Series Classification
by: Cai, Huili, et al.
Published: (2023)
by: Cai, Huili, et al.
Published: (2023)
End-to-end Deep Reinforcement Learning for Stochastic Multi-objective Optimization in C-VRPTW
by: Abouelrous, Abdo, et al.
Published: (2025)
by: Abouelrous, Abdo, et al.
Published: (2025)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
by: Zheng, Yinan, et al.
Published: (2026)
by: Zheng, Yinan, et al.
Published: (2026)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Similar Items
-
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026) -
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025) -
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026) -
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026) -
Simple, unified analysis of Johnson-Lindenstrauss with applications
by: Li, Yingru
Published: (2024)