Scalable Ensembling For Mitigating Reward Overoptimisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmed, Ahmed M., Rafailov, Rafael, Sharkov, Stepan, Li, Xuechen, Koyejo, Sanmi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916293795381248
author Ahmed, Ahmed M.
Rafailov, Rafael
Sharkov, Stepan
Li, Xuechen
Koyejo, Sanmi
author_facet Ahmed, Ahmed M.
Rafailov, Rafael
Sharkov, Stepan
Li, Xuechen
Koyejo, Sanmi
contents Reinforcement Learning from Human Feedback (RLHF) has enabled significant advancements within language modeling for powerful, instruction-following models. However, the alignment of these models remains a pressing challenge as the policy tends to overfit the learned ``proxy" reward model past an inflection point of utility as measured by a ``gold" reward model that is more performant -- a phenomenon known as overoptimisation. Prior work has mitigated this issue by computing a pessimistic statistic over an ensemble of reward models, which is common in Offline Reinforcement Learning but incredibly costly for language models with high memory requirements, making such approaches infeasible for sufficiently large models. To this end, we propose using a shared encoder but separate linear heads. We find this leads to similar performance as the full ensemble while allowing tremendous savings in memory and time required for training for models of similar size.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01013
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable Ensembling For Mitigating Reward Overoptimisation
Ahmed, Ahmed M.
Rafailov, Rafael
Sharkov, Stepan
Li, Xuechen
Koyejo, Sanmi
Machine Learning
Computation and Language
Reinforcement Learning from Human Feedback (RLHF) has enabled significant advancements within language modeling for powerful, instruction-following models. However, the alignment of these models remains a pressing challenge as the policy tends to overfit the learned ``proxy" reward model past an inflection point of utility as measured by a ``gold" reward model that is more performant -- a phenomenon known as overoptimisation. Prior work has mitigated this issue by computing a pessimistic statistic over an ensemble of reward models, which is common in Offline Reinforcement Learning but incredibly costly for language models with high memory requirements, making such approaches infeasible for sufficiently large models. To this end, we propose using a shared encoder but separate linear heads. We find this leads to similar performance as the full ensemble while allowing tremendous savings in memory and time required for training for models of similar size.
title Scalable Ensembling For Mitigating Reward Overoptimisation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2406.01013