Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Yifeng, Le, Hung, Han, Songyang, Ruan, Kangrui, Jin, Zhenghui, Kumar, Varun, Wang, Zijian, Deoras, Anoop
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913046740336640
author Ding, Yifeng
Le, Hung
Han, Songyang
Ruan, Kangrui
Jin, Zhenghui
Kumar, Varun
Wang, Zijian
Deoras, Anoop
author_facet Ding, Yifeng
Le, Hung
Han, Songyang
Ruan, Kangrui
Jin, Zhenghui
Kumar, Varun
Wang, Zijian
Deoras, Anoop
contents Training Large Language Models (LLMs) for multi-turn Tool-Integrated Reasoning (TIR) - where models iteratively reason, generate code, and verify through execution - remains challenging for existing reinforcement learning (RL) approaches. Current RL methods, exemplified by Group Relative Policy Optimization (GRPO), suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation. To address this issue, we propose Group Turn Policy Optimization (GTPO), a novel RL algorithm specifically designed for training LLMs on multi-turn TIR tasks. GTPO introduces three key innovations: (1) turn-level reward assignment that provides fine-grained feedback for individual turns, (2) return-based advantage estimation where normalized discounted returns are calculated as advantages, and (3) self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards. Our comprehensive evaluation demonstrates that GTPO outperforms GRPO by 3.0% across diverse math reasoning benchmarks, establishing its effectiveness. GTPO also improves GRPO by 3.9% on commonsense reasoning and program synthesis tasks, demonstrating its generalizability to non-math domains. Importantly, GTPO incurs negligible overhead, ensuring its practicality for real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
Ding, Yifeng
Le, Hung
Han, Songyang
Ruan, Kangrui
Jin, Zhenghui
Kumar, Varun
Wang, Zijian
Deoras, Anoop
Machine Learning
Artificial Intelligence
Computation and Language
Training Large Language Models (LLMs) for multi-turn Tool-Integrated Reasoning (TIR) - where models iteratively reason, generate code, and verify through execution - remains challenging for existing reinforcement learning (RL) approaches. Current RL methods, exemplified by Group Relative Policy Optimization (GRPO), suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation. To address this issue, we propose Group Turn Policy Optimization (GTPO), a novel RL algorithm specifically designed for training LLMs on multi-turn TIR tasks. GTPO introduces three key innovations: (1) turn-level reward assignment that provides fine-grained feedback for individual turns, (2) return-based advantage estimation where normalized discounted returns are calculated as advantages, and (3) self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards. Our comprehensive evaluation demonstrates that GTPO outperforms GRPO by 3.0% across diverse math reasoning benchmarks, establishing its effectiveness. GTPO also improves GRPO by 3.9% on commonsense reasoning and program synthesis tasks, demonstrating its generalizability to non-math domains. Importantly, GTPO incurs negligible overhead, ensuring its practicality for real-world scenarios.
title Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.14846