Improving Reward Models with Synthetic Critiques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Zihuiwen, Greenlee-Scott, Fraser, Bartolo, Max, Blunsom, Phil, Campos, Jon Ander, Gallé, Matthias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916443982921728
author Ye, Zihuiwen
Greenlee-Scott, Fraser
Bartolo, Max
Blunsom, Phil
Campos, Jon Ander
Gallé, Matthias
author_facet Ye, Zihuiwen
Greenlee-Scott, Fraser
Bartolo, Max
Blunsom, Phil
Campos, Jon Ander
Gallé, Matthias
contents Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback. RMs are trained to predict a score reflecting human preference, which requires significant time and cost for human annotation. Additionally, RMs tend to quickly overfit on superficial features in the training set, hindering their generalization performance on unseen distributions. We propose a novel approach using synthetic natural language critiques generated by large language models to provide additional feedback, evaluating aspects such as instruction following, correctness, and style. This offers richer signals and more robust features for RMs to assess and score on. We demonstrate that high-quality critiques improve the performance and data efficiency of RMs initialized from different pretrained models, reducing the reliance on costly human annotations. Furthermore, incorporating critiques improves both the interpretability and robustness of RM training.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20850
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Reward Models with Synthetic Critiques
Ye, Zihuiwen
Greenlee-Scott, Fraser
Bartolo, Max
Blunsom, Phil
Campos, Jon Ander
Gallé, Matthias
Computation and Language
Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback. RMs are trained to predict a score reflecting human preference, which requires significant time and cost for human annotation. Additionally, RMs tend to quickly overfit on superficial features in the training set, hindering their generalization performance on unseen distributions. We propose a novel approach using synthetic natural language critiques generated by large language models to provide additional feedback, evaluating aspects such as instruction following, correctness, and style. This offers richer signals and more robust features for RMs to assess and score on. We demonstrate that high-quality critiques improve the performance and data efficiency of RMs initialized from different pretrained models, reducing the reliance on costly human annotations. Furthermore, incorporating critiques improves both the interpretability and robustness of RM training.
title Improving Reward Models with Synthetic Critiques
topic Computation and Language
url https://arxiv.org/abs/2405.20850