Guardado en:
Detalles Bibliográficos
Autores principales: zhang, Ranxu, li, zeyang, Huang, Jiacheng, Zhang, Rui, Xu, Xiaozhou, zhe, sun, Zhang, Yanyong, Wang, Chao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2605.23382
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916038608683008
author zhang, Ranxu
li, zeyang
Huang, Jiacheng
Zhang, Rui
Xu, Xiaozhou
zhe, sun
Zhang, Yanyong
Wang, Chao
author_facet zhang, Ranxu
li, zeyang
Huang, Jiacheng
Zhang, Rui
Xu, Xiaozhou
zhe, sun
Zhang, Yanyong
Wang, Chao
contents Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned behavior: the same query may call for different planning strategies and tool-use decisions across users. This setting raises key challenges: generic rewards cannot capture heterogeneous user preferences, observed behaviors are entangled with conformity effects, and flat memories cannot support personalized skill retrieval. To this end, we propose a unified personalized Agentic RL framework that embeds personalization into training-time optimization. At its core is \emph{Personalized Anchor Reward-Decoupled Policy Optimization} (\textbf{PARPO}), which decouples generic task-quality rewards from personalized preference rewards and uses user-specific anchors to stabilize learning under heterogeneous reward scales. We further introduce a two-stage preference-disentangled reward model and \emph{Preference-Aligned Skill Evolution Graph Memory} (\textbf{PSGM}) for personalized supervision and preference-aligned skill retrieval. Together, they form a closed loop of preference identification, policy optimization, and structured skill accumulation. Experiments on ETAPP, ETAPP-Hard, and SJAgent show that our framework consistently outperforms strong memory and RL baselines. Code and data are included in the supplementary materials.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23382
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
zhang, Ranxu
li, zeyang
Huang, Jiacheng
Zhang, Rui
Xu, Xiaozhou
zhe, sun
Zhang, Yanyong
Wang, Chao
Computation and Language
Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned behavior: the same query may call for different planning strategies and tool-use decisions across users. This setting raises key challenges: generic rewards cannot capture heterogeneous user preferences, observed behaviors are entangled with conformity effects, and flat memories cannot support personalized skill retrieval. To this end, we propose a unified personalized Agentic RL framework that embeds personalization into training-time optimization. At its core is \emph{Personalized Anchor Reward-Decoupled Policy Optimization} (\textbf{PARPO}), which decouples generic task-quality rewards from personalized preference rewards and uses user-specific anchors to stabilize learning under heterogeneous reward scales. We further introduce a two-stage preference-disentangled reward model and \emph{Preference-Aligned Skill Evolution Graph Memory} (\textbf{PSGM}) for personalized supervision and preference-aligned skill retrieval. Together, they form a closed loop of preference identification, policy optimization, and structured skill accumulation. Experiments on ETAPP, ETAPP-Hard, and SJAgent show that our framework consistently outperforms strong memory and RL baselines. Code and data are included in the supplementary materials.
title From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2605.23382