PILAF: Optimal Human Preference Sampling for Reward Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Feng, Yunzhen, Kwiatkowski, Ariel, Zheng, Kunhao, Kempe, Julia, Duan, Yaqi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912222517657600
author Feng, Yunzhen
Kwiatkowski, Ariel
Zheng, Kunhao
Kempe, Julia
Duan, Yaqi
author_facet Feng, Yunzhen
Kwiatkowski, Ariel
Zheng, Kunhao
Kempe, Julia
Duan, Yaqi
contents As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating preference data into reward models when oracle human values remain inaccessible. In practice, RLHF mostly relies on approximate reward models, which may not consistently guide the policy toward maximizing the underlying human values. We propose Policy-Interpolated Learning for Aligned Feedback (PILAF), a novel response sampling strategy for preference labeling that explicitly aligns preference learning with maximizing the underlying oracle reward. PILAF is theoretically grounded, demonstrating optimality from both an optimization and a statistical perspective. The method is straightforward to implement and demonstrates strong performance in iterative and online RLHF settings where feedback curation is critical.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PILAF: Optimal Human Preference Sampling for Reward Modeling
Feng, Yunzhen
Kwiatkowski, Ariel
Zheng, Kunhao
Kempe, Julia
Duan, Yaqi
Machine Learning
As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating preference data into reward models when oracle human values remain inaccessible. In practice, RLHF mostly relies on approximate reward models, which may not consistently guide the policy toward maximizing the underlying human values. We propose Policy-Interpolated Learning for Aligned Feedback (PILAF), a novel response sampling strategy for preference labeling that explicitly aligns preference learning with maximizing the underlying oracle reward. PILAF is theoretically grounded, demonstrating optimality from both an optimization and a statistical perspective. The method is straightforward to implement and demonstrates strong performance in iterative and online RLHF settings where feedback curation is critical.
title PILAF: Optimal Human Preference Sampling for Reward Modeling
topic Machine Learning
url https://arxiv.org/abs/2502.04270