Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Rui, Ding, Ruomeng, Lin, Yong, Zhang, Huan, Zhang, Tong
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912082707873792
author Yang, Rui
Ding, Ruomeng
Lin, Yong
Zhang, Huan
Zhang, Tong
author_facet Yang, Rui
Ding, Ruomeng
Lin, Yong
Zhang, Huan
Zhang, Tong
contents Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning from human feedback (RLHF). However, current reward models have limited generalization capabilities to unseen prompts and responses, which can lead to an unexpected phenomenon known as reward over-optimization, resulting in a decline in actual performance due to excessive optimization of rewards. While previous research has advocated for constraining policy optimization, our study introduces a novel approach to enhance the reward model's generalization ability against distribution shifts by regularizing the hidden states. Specifically, we retain the base model's language model head and incorporate a suite of text-generation losses to preserve the hidden states' text-generation capabilities, while concurrently learning a reward head behind the same hidden states. Our experimental results demonstrate that the introduced regularization technique markedly improves the accuracy of learned reward models across a variety of out-of-distribution (OOD) tasks and effectively alleviates the over-optimization issue in RLHF, offering a more reliable and robust preference learning paradigm.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10216
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Yang, Rui
Ding, Ruomeng
Lin, Yong
Zhang, Huan
Zhang, Tong
Computation and Language
Artificial Intelligence
Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning from human feedback (RLHF). However, current reward models have limited generalization capabilities to unseen prompts and responses, which can lead to an unexpected phenomenon known as reward over-optimization, resulting in a decline in actual performance due to excessive optimization of rewards. While previous research has advocated for constraining policy optimization, our study introduces a novel approach to enhance the reward model's generalization ability against distribution shifts by regularizing the hidden states. Specifically, we retain the base model's language model head and incorporate a suite of text-generation losses to preserve the hidden states' text-generation capabilities, while concurrently learning a reward head behind the same hidden states. Our experimental results demonstrate that the introduced regularization technique markedly improves the accuracy of learned reward models across a variety of out-of-distribution (OOD) tasks and effectively alleviates the over-optimization issue in RLHF, offering a more reliable and robust preference learning paradigm.
title Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.10216