Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haoyu, Chen, Yuxin, Luo, Liang, Zhang, Buyun, Wen, Ellie Dingqiao, Li, Pan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911542932406272
author Wang, Haoyu
Chen, Yuxin
Luo, Liang
Zhang, Buyun
Wen, Ellie Dingqiao
Li, Pan
author_facet Wang, Haoyu
Chen, Yuxin
Luo, Liang
Zhang, Buyun
Wen, Ellie Dingqiao
Li, Pan
contents Multi-turn human-AI collaboration is fundamental to deploying interactive services such as adaptive tutoring, conversational recommendation, and professional consultation. However, optimizing these interactions via reinforcement learning is hindered by the sparsity of verifiable intermediate rewards and the high stochasticity of user responses. To address these challenges, we introduce Implicit Turn-wise Policy Optimization (ITPO). ITPO leverages an implicit process reward model to derive fine-grained, turn-wise process rewards from sparse outcome signals. Unlike volatile token-level rewards, these turn-level signals exhibit superior robustness and may utilize a normalization mechanism to further enhance training stability. We evaluate ITPO across three representative multi-turn collaborative tasks: math tutoring, document writing, and medical recommendation. Empirical results demonstrate that ITPO, when combined with PPO, GRPO, or RLOO, consistently achieves improved convergence than existing baselines. Elaborate trajectory analysis confirms that ITPO infers turn-wise preferences that are semantically aligned with human judgment. Code is publicly available at https://github.com/Graph-COM/ITPO.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23550
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
Wang, Haoyu
Chen, Yuxin
Luo, Liang
Zhang, Buyun
Wen, Ellie Dingqiao
Li, Pan
Machine Learning
Multi-turn human-AI collaboration is fundamental to deploying interactive services such as adaptive tutoring, conversational recommendation, and professional consultation. However, optimizing these interactions via reinforcement learning is hindered by the sparsity of verifiable intermediate rewards and the high stochasticity of user responses. To address these challenges, we introduce Implicit Turn-wise Policy Optimization (ITPO). ITPO leverages an implicit process reward model to derive fine-grained, turn-wise process rewards from sparse outcome signals. Unlike volatile token-level rewards, these turn-level signals exhibit superior robustness and may utilize a normalization mechanism to further enhance training stability. We evaluate ITPO across three representative multi-turn collaborative tasks: math tutoring, document writing, and medical recommendation. Empirical results demonstrate that ITPO, when combined with PPO, GRPO, or RLOO, consistently achieves improved convergence than existing baselines. Elaborate trajectory analysis confirms that ITPO infers turn-wise preferences that are semantically aligned with human judgment. Code is publicly available at https://github.com/Graph-COM/ITPO.
title Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
topic Machine Learning
url https://arxiv.org/abs/2603.23550