Information-Theoretic Reward Decomposition for Generalizable RLHF

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mao, Liyuan, Xu, Haoran, Zhang, Amy, Zhang, Weinan, Bai, Chenjia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915572806057984
author Mao, Liyuan
Xu, Haoran
Zhang, Amy
Zhang, Weinan
Bai, Chenjia
author_facet Mao, Liyuan
Xu, Haoran
Zhang, Amy
Zhang, Weinan
Bai, Chenjia
contents A generalizable reward model is crucial in Reinforcement Learning from Human Feedback (RLHF) as it enables correctly evaluating unseen prompt-response pairs. However, existing reward models lack this ability, as they are typically trained by increasing the reward gap between chosen and rejected responses, while overlooking the prompts that the responses are conditioned on. Consequently, when the trained reward model is evaluated on prompt-response pairs that lie outside the data distribution, neglecting the effect of prompts may result in poor generalization of the reward model. To address this issue, we decompose the reward value into two independent components: prompt-free reward and prompt-related reward. Prompt-free reward represents the evaluation that is determined only by responses, while the prompt-related reward reflects the reward that derives from both the prompt and the response. We extract these two components from an information-theoretic perspective, which requires no extra models. Subsequently, we propose a new reward learning algorithm by prioritizing data samples based on their prompt-free reward values. Through toy examples, we demonstrate that the extracted prompt-free and prompt-related rewards effectively characterize two parts of the reward model. Further, standard evaluations show that our method improves both the alignment performance and the generalization capability of the reward model.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06020
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Information-Theoretic Reward Decomposition for Generalizable RLHF
Mao, Liyuan
Xu, Haoran
Zhang, Amy
Zhang, Weinan
Bai, Chenjia
Artificial Intelligence
Computation and Language
Machine Learning
A generalizable reward model is crucial in Reinforcement Learning from Human Feedback (RLHF) as it enables correctly evaluating unseen prompt-response pairs. However, existing reward models lack this ability, as they are typically trained by increasing the reward gap between chosen and rejected responses, while overlooking the prompts that the responses are conditioned on. Consequently, when the trained reward model is evaluated on prompt-response pairs that lie outside the data distribution, neglecting the effect of prompts may result in poor generalization of the reward model. To address this issue, we decompose the reward value into two independent components: prompt-free reward and prompt-related reward. Prompt-free reward represents the evaluation that is determined only by responses, while the prompt-related reward reflects the reward that derives from both the prompt and the response. We extract these two components from an information-theoretic perspective, which requires no extra models. Subsequently, we propose a new reward learning algorithm by prioritizing data samples based on their prompt-free reward values. Through toy examples, we demonstrate that the extracted prompt-free and prompt-related rewards effectively characterize two parts of the reward model. Further, standard evaluations show that our method improves both the alignment performance and the generalization capability of the reward model.
title Information-Theoretic Reward Decomposition for Generalizable RLHF
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.06020