Policy-labeled Preference Learning: Is Preference Enough for RLHF?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cho, Taehyun, Ju, Seokhun, Han, Seungyub, Kim, Dohyeong, Lee, Kyungjae, Lee, Jungwoo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915285764669440
author Cho, Taehyun
Ju, Seokhun
Han, Seungyub
Kim, Dohyeong
Lee, Kyungjae
Lee, Jungwoo
author_facet Cho, Taehyun
Ju, Seokhun
Han, Seungyub
Kim, Dohyeong
Lee, Kyungjae
Lee, Jungwoo
contents To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing policies via reinforcement learning algorithms. However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. Inspired by Direct Preference Optimization framework which directly learns optimal policy without explicit reward, we propose policy-labeled preference learning (PPL), to resolve likelihood mismatch issues by modeling human preferences with regret, which reflects behavior policy information. We also provide a contrastive KL regularization, derived from regret-based principles, to enhance RLHF in sequential decision making. Experiments in high-dimensional continuous control tasks demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy-labeled Preference Learning: Is Preference Enough for RLHF?
Cho, Taehyun
Ju, Seokhun
Han, Seungyub
Kim, Dohyeong
Lee, Kyungjae
Lee, Jungwoo
Machine Learning
Artificial Intelligence
To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing policies via reinforcement learning algorithms. However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. Inspired by Direct Preference Optimization framework which directly learns optimal policy without explicit reward, we propose policy-labeled preference learning (PPL), to resolve likelihood mismatch issues by modeling human preferences with regret, which reflects behavior policy information. We also provide a contrastive KL regularization, derived from regret-based principles, to enhance RLHF in sequential decision making. Experiments in high-dimensional continuous control tasks demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings.
title Policy-labeled Preference Learning: Is Preference Enough for RLHF?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.06273