A Long Way to Go: Investigating Length Correlations in RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singhal, Prasann, Goyal, Tanya, Xu, Jiacheng, Durrett, Greg
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917718575284224
author Singhal, Prasann
Goyal, Tanya
Xu, Jiacheng
Durrett, Greg
author_facet Singhal, Prasann
Goyal, Tanya
Xu, Jiacheng
Durrett, Greg
contents Great success has been reported using Reinforcement Learning from Human Feedback (RLHF) to align large language models, with open preference datasets enabling wider experimentation, particularly for "helpfulness" in tasks like dialogue and web question answering. Alongside these improvements, however, RLHF also often drives models to produce longer outputs. This paper demonstrates, on three diverse settings, that optimizing for response length is, much more than previously thought, a significant factor behind RLHF. Studying the strategies RL optimization uses to maximize reward, we find improvements in reward to largely be driven by increasing response length, instead of other features. Indeed, we find that even a purely length-based reward reproduces most downstream RLHF improvements over supervised fine-tuned models. Testing a comprehensive set of length-countering interventions, we identify the dominant source of these biases to be reward models, which, by studying training dynamics, we find are non-robust and easily influenced by length biases in preference data.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03716
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Long Way to Go: Investigating Length Correlations in RLHF
Singhal, Prasann
Goyal, Tanya
Xu, Jiacheng
Durrett, Greg
Computation and Language
Machine Learning
Great success has been reported using Reinforcement Learning from Human Feedback (RLHF) to align large language models, with open preference datasets enabling wider experimentation, particularly for "helpfulness" in tasks like dialogue and web question answering. Alongside these improvements, however, RLHF also often drives models to produce longer outputs. This paper demonstrates, on three diverse settings, that optimizing for response length is, much more than previously thought, a significant factor behind RLHF. Studying the strategies RL optimization uses to maximize reward, we find improvements in reward to largely be driven by increasing response length, instead of other features. Indeed, we find that even a purely length-based reward reproduces most downstream RLHF improvements over supervised fine-tuned models. Testing a comprehensive set of length-countering interventions, we identify the dominant source of these biases to be reward models, which, by studying training dynamics, we find are non-robust and easily influenced by length biases in preference data.
title A Long Way to Go: Investigating Length Correlations in RLHF
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.03716