Saved in:
Bibliographic Details
Main Authors: Li, Junbo, Zhou, Peng, Meng, Rui, Vadera, Meet P., Li, Lihong, Li, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.17008
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909999628812288
author Li, Junbo
Zhou, Peng
Meng, Rui
Vadera, Meet P.
Li, Lihong
Li, Yang
author_facet Li, Junbo
Zhou, Peng
Meng, Rui
Vadera, Meet P.
Li, Lihong
Li, Yang
contents Reinforcement learning (RL) has re-emerged as a natural approach for training interactive LLM agents in real-world environments. However, directly applying the widely used Group Relative Policy Optimization (GRPO) algorithm to multi-turn tasks exposes notable limitations, particularly in scenarios requiring long-horizon reasoning. To address these challenges, we investigate more stable and effective advantage estimation strategies, especially for multi-turn settings. We first explore Proximal Policy Optimization (PPO) as an alternative and find it to be more robust than GRPO. To further enhance PPO in multi-turn scenarios, we introduce turn-PPO, a variant that operates on a turn-level MDP formulation, as opposed to the commonly used token-level MDP. Our results on the WebShop and Sokoban datasets demonstrate the effectiveness of turn-PPO, both with and without long reasoning components.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
Li, Junbo
Zhou, Peng
Meng, Rui
Vadera, Meet P.
Li, Lihong
Li, Yang
Machine Learning
Reinforcement learning (RL) has re-emerged as a natural approach for training interactive LLM agents in real-world environments. However, directly applying the widely used Group Relative Policy Optimization (GRPO) algorithm to multi-turn tasks exposes notable limitations, particularly in scenarios requiring long-horizon reasoning. To address these challenges, we investigate more stable and effective advantage estimation strategies, especially for multi-turn settings. We first explore Proximal Policy Optimization (PPO) as an alternative and find it to be more robust than GRPO. To further enhance PPO in multi-turn scenarios, we introduce turn-PPO, a variant that operates on a turn-level MDP formulation, as opposed to the commonly used token-level MDP. Our results on the WebShop and Sokoban datasets demonstrate the effectiveness of turn-PPO, both with and without long reasoning components.
title Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
topic Machine Learning
url https://arxiv.org/abs/2512.17008