COPR: Continual Human Preference Learning via Optimal Policy Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Han, Gui, Lin, Lei, Yu, Zhai, Yuanzhao, Zhang, Yehong, He, Yulan, Wang, Hui, Yu, Yue, Wong, Kam-Fai, Liang, Bin, Xu, Ruifeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929642907107328
author Zhang, Han
Gui, Lin
Lei, Yu
Zhai, Yuanzhao
Zhang, Yehong
He, Yulan
Wang, Hui
Yu, Yue
Wong, Kam-Fai
Liang, Bin
Xu, Ruifeng
author_facet Zhang, Han
Gui, Lin
Lei, Yu
Zhai, Yuanzhao
Zhang, Yehong
He, Yulan
Wang, Hui
Yu, Yue
Wong, Kam-Fai
Liang, Bin
Xu, Ruifeng
contents Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical in comparison to traditional static alignment. Nevertheless, making RLHF compatible with Continual Learning (CL) is challenging due to its complex process. Meanwhile, directly learning new human preferences may lead to Catastrophic Forgetting (CF) of historical preferences, resulting in helpless or harmful outputs. To overcome these challenges, we propose the Continual Optimal Policy Regularization (COPR) method, which draws inspiration from the optimal policy theory. COPR utilizes a sampling distribution as a demonstration and regularization constraints for CL. It adopts the Lagrangian Duality (LD) method to dynamically regularize the current policy based on the historically optimal policy, which prevents CF and avoids over-emphasizing unbalanced objectives. We also provide formal proof for the learnability of COPR. The experimental results show that COPR outperforms strong CL baselines on our proposed benchmark, in terms of reward-based, GPT-4 evaluations and human assessment. Furthermore, we validate the robustness of COPR under various CL settings, including different backbones, replay memory sizes, and learning orders.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14228
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COPR: Continual Human Preference Learning via Optimal Policy Regularization
Zhang, Han
Gui, Lin
Lei, Yu
Zhai, Yuanzhao
Zhang, Yehong
He, Yulan
Wang, Hui
Yu, Yue
Wong, Kam-Fai
Liang, Bin
Xu, Ruifeng
Machine Learning
Artificial Intelligence
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical in comparison to traditional static alignment. Nevertheless, making RLHF compatible with Continual Learning (CL) is challenging due to its complex process. Meanwhile, directly learning new human preferences may lead to Catastrophic Forgetting (CF) of historical preferences, resulting in helpless or harmful outputs. To overcome these challenges, we propose the Continual Optimal Policy Regularization (COPR) method, which draws inspiration from the optimal policy theory. COPR utilizes a sampling distribution as a demonstration and regularization constraints for CL. It adopts the Lagrangian Duality (LD) method to dynamically regularize the current policy based on the historically optimal policy, which prevents CF and avoids over-emphasizing unbalanced objectives. We also provide formal proof for the learnability of COPR. The experimental results show that COPR outperforms strong CL baselines on our proposed benchmark, in terms of reward-based, GPT-4 evaluations and human assessment. Furthermore, we validate the robustness of COPR under various CL settings, including different backbones, replay memory sizes, and learning orders.
title COPR: Continual Human Preference Learning via Optimal Policy Regularization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.14228