WPO: Enhancing RLHF with Weighted Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Wenxuan, Agrawal, Ravi, Zhang, Shujian, Indurthi, Sathish Reddy, Zhao, Sanqiang, Song, Kaiqiang, Xu, Silei, Zhu, Chenguang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929526233104384
author Zhou, Wenxuan
Agrawal, Ravi
Zhang, Shujian
Indurthi, Sathish Reddy
Zhao, Sanqiang
Song, Kaiqiang
Xu, Silei
Zhu, Chenguang
author_facet Zhou, Wenxuan
Agrawal, Ravi
Zhang, Shujian
Indurthi, Sathish Reddy
Zhao, Sanqiang
Song, Kaiqiang
Xu, Silei
Zhu, Chenguang
contents Reinforcement learning from human feedback (RLHF) is a promising solution to align large language models (LLMs) more closely with human values. Off-policy preference optimization, where the preference data is obtained from other models, is widely adopted due to its cost efficiency and scalability. However, off-policy preference optimization often suffers from a distributional gap between the policy used for data collection and the target policy, leading to suboptimal optimization. In this paper, we propose a novel strategy to mitigate this problem by simulating on-policy learning with off-policy preference data. Our Weighted Preference Optimization (WPO) method adapts off-policy data to resemble on-policy data more closely by reweighting preference pairs according to their probability under the current policy. This method not only addresses the distributional gap problem but also enhances the optimization process without incurring additional costs. We validate our method on instruction following benchmarks including Alpaca Eval 2 and MT-bench. WPO not only outperforms Direct Preference Optimization (DPO) by up to 5.6% on Alpaca Eval 2 but also establishes a remarkable length-controlled winning rate against GPT-4-turbo of 76.7% based on Gemma-2-9b-it. We release the code and models at https://github.com/wzhouad/WPO.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11827
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WPO: Enhancing RLHF with Weighted Preference Optimization
Zhou, Wenxuan
Agrawal, Ravi
Zhang, Shujian
Indurthi, Sathish Reddy
Zhao, Sanqiang
Song, Kaiqiang
Xu, Silei
Zhu, Chenguang
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement learning from human feedback (RLHF) is a promising solution to align large language models (LLMs) more closely with human values. Off-policy preference optimization, where the preference data is obtained from other models, is widely adopted due to its cost efficiency and scalability. However, off-policy preference optimization often suffers from a distributional gap between the policy used for data collection and the target policy, leading to suboptimal optimization. In this paper, we propose a novel strategy to mitigate this problem by simulating on-policy learning with off-policy preference data. Our Weighted Preference Optimization (WPO) method adapts off-policy data to resemble on-policy data more closely by reweighting preference pairs according to their probability under the current policy. This method not only addresses the distributional gap problem but also enhances the optimization process without incurring additional costs. We validate our method on instruction following benchmarks including Alpaca Eval 2 and MT-bench. WPO not only outperforms Direct Preference Optimization (DPO) by up to 5.6% on Alpaca Eval 2 but also establishes a remarkable length-controlled winning rate against GPT-4-turbo of 76.7% based on Gemma-2-9b-it. We release the code and models at https://github.com/wzhouad/WPO.
title WPO: Enhancing RLHF with Weighted Preference Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.11827