Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Weixiang, Sui, Xingyu, Hu, Yulin, Guo, Jiahe, Liu, Haixiao, Li, Biye, Zhao, Yanyan, Qin, Bing, Liu, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908704622772224
author Zhao, Weixiang
Sui, Xingyu
Hu, Yulin
Guo, Jiahe
Liu, Haixiao
Li, Biye
Zhao, Yanyan
Qin, Bing
Liu, Ting
author_facet Zhao, Weixiang
Sui, Xingyu
Hu, Yulin
Guo, Jiahe
Liu, Haixiao
Li, Biye
Zhao, Yanyan
Qin, Bing
Liu, Ting
contents Personalized alignment is essential for enabling large language models (LLMs) to engage effectively in user-centric dialogue. While recent prompt-based and offline optimization methods offer preliminary solutions, they fall short in cold-start scenarios and long-term personalization due to their inherently static and shallow designs. In this work, we introduce the Reinforcement Learning for Personalized Alignment (RLPA) framework, in which an LLM interacts with a simulated user model to iteratively infer and refine user profiles through dialogue. The training process is guided by a dual-level reward structure: the Profile Reward encourages accurate construction of user representations, while the Response Reward incentivizes generation of responses consistent with the inferred profile. We instantiate RLPA by fine-tuning Qwen-2.5-3B-Instruct, resulting in Qwen-RLPA, which achieves state-of-the-art performance in personalized dialogue. Empirical evaluations demonstrate that Qwen-RLPA consistently outperforms prompting and offline fine-tuning baselines, and even surpasses advanced commercial models such as Claude-3.5 and GPT-4o. Further analysis highlights Qwen-RLPA's robustness in reconciling conflicting user preferences, sustaining long-term personalization and delivering more efficient inference compared to recent reasoning-focused LLMs. These results emphasize the potential of dynamic profile inference as a more effective paradigm for building personalized dialogue systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
Zhao, Weixiang
Sui, Xingyu
Hu, Yulin
Guo, Jiahe
Liu, Haixiao
Li, Biye
Zhao, Yanyan
Qin, Bing
Liu, Ting
Computation and Language
Personalized alignment is essential for enabling large language models (LLMs) to engage effectively in user-centric dialogue. While recent prompt-based and offline optimization methods offer preliminary solutions, they fall short in cold-start scenarios and long-term personalization due to their inherently static and shallow designs. In this work, we introduce the Reinforcement Learning for Personalized Alignment (RLPA) framework, in which an LLM interacts with a simulated user model to iteratively infer and refine user profiles through dialogue. The training process is guided by a dual-level reward structure: the Profile Reward encourages accurate construction of user representations, while the Response Reward incentivizes generation of responses consistent with the inferred profile. We instantiate RLPA by fine-tuning Qwen-2.5-3B-Instruct, resulting in Qwen-RLPA, which achieves state-of-the-art performance in personalized dialogue. Empirical evaluations demonstrate that Qwen-RLPA consistently outperforms prompting and offline fine-tuning baselines, and even surpasses advanced commercial models such as Claude-3.5 and GPT-4o. Further analysis highlights Qwen-RLPA's robustness in reconciling conflicting user preferences, sustaining long-term personalization and delivering more efficient inference compared to recent reasoning-focused LLMs. These results emphasize the potential of dynamic profile inference as a more effective paradigm for building personalized dialogue systems.
title Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
topic Computation and Language
url https://arxiv.org/abs/2505.15456