Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rita, Mathieu, Strub, Florian, Chaabouni, Rahma, Michel, Paul, Dupoux, Emmanuel, Pietquin, Olivier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911860068974592
author Rita, Mathieu
Strub, Florian
Chaabouni, Rahma
Michel, Paul
Dupoux, Emmanuel
Pietquin, Olivier
author_facet Rita, Mathieu
Strub, Florian
Chaabouni, Rahma
Michel, Paul
Dupoux, Emmanuel
Pietquin, Olivier
contents While Reinforcement Learning (RL) has been proven essential for tuning large language models (LLMs), it can lead to reward over-optimization (ROO). Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. Additionally, KL regularization focuses solely on regularizing the language policy, neglecting a potential source of regularization: the reward function itself. Inspired by demonstration-guided RL, we here introduce the Reward Calibration from Demonstration (RCfD), which leverages human demonstrations and a reward model to recalibrate the reward objective. Formally, given a prompt, the RCfD objective minimizes the distance between the demonstrations' and LLM's rewards rather than directly maximizing the reward function. This objective shift avoids incentivizing the LLM to exploit the reward model and promotes more natural and diverse language generation. We show the effectiveness of RCfD on three language tasks, which achieves comparable performance to carefully tuned baselines while mitigating ROO.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19409
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
Rita, Mathieu
Strub, Florian
Chaabouni, Rahma
Michel, Paul
Dupoux, Emmanuel
Pietquin, Olivier
Computation and Language
While Reinforcement Learning (RL) has been proven essential for tuning large language models (LLMs), it can lead to reward over-optimization (ROO). Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. Additionally, KL regularization focuses solely on regularizing the language policy, neglecting a potential source of regularization: the reward function itself. Inspired by demonstration-guided RL, we here introduce the Reward Calibration from Demonstration (RCfD), which leverages human demonstrations and a reward model to recalibrate the reward objective. Formally, given a prompt, the RCfD objective minimizes the distance between the demonstrations' and LLM's rewards rather than directly maximizing the reward function. This objective shift avoids incentivizing the LLM to exploit the reward model and promotes more natural and diverse language generation. We show the effectiveness of RCfD on three language tasks, which achieves comparable performance to carefully tuned baselines while mitigating ROO.
title Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2404.19409