Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Zhenyu, Du, Pengfan, Niu, Yilin, Du, Zhengxiao, Zeng, Aohan, Liu, Xiao, Huang, Minlie, Wang, Hongning, Tang, Jie, Dong, Yuxiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917861563301888
author Hou, Zhenyu
Du, Pengfan
Niu, Yilin
Du, Zhengxiao
Zeng, Aohan
Liu, Xiao
Huang, Minlie
Wang, Hongning
Tang, Jie
Dong, Yuxiao
author_facet Hou, Zhenyu
Du, Pengfan
Niu, Yilin
Du, Zhengxiao
Zeng, Aohan
Liu, Xiao
Huang, Minlie
Wang, Hongning
Tang, Jie
Dong, Yuxiao
contents This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs). Although RLHF is considered an important step in post-training of LLMs, its scaling potential is still largely unknown. We systematically analyze key components in the RLHF framework--model size, data composition, and inference budget--and their impacts on performance. Our findings show that increasing data diversity and volume improves reward model performance, helping process-supervision models scale better. For policy training, more response samples per prompt boost performance initially but quickly plateau. And larger reward models offer modest gains in policy training. In addition, larger policy models benefit less from RLHF with a fixed reward model. Overall, RLHF scales less efficiently than pretraining, with diminishing returns from additional computational resources. Based on these observations, we propose strategies to optimize RLHF performance within computational limits.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06000
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
Hou, Zhenyu
Du, Pengfan
Niu, Yilin
Du, Zhengxiao
Zeng, Aohan
Liu, Xiao
Huang, Minlie
Wang, Hongning
Tang, Jie
Dong, Yuxiao
Computation and Language
Machine Learning
This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs). Although RLHF is considered an important step in post-training of LLMs, its scaling potential is still largely unknown. We systematically analyze key components in the RLHF framework--model size, data composition, and inference budget--and their impacts on performance. Our findings show that increasing data diversity and volume improves reward model performance, helping process-supervision models scale better. For policy training, more response samples per prompt boost performance initially but quickly plateau. And larger reward models offer modest gains in policy training. In addition, larger policy models benefit less from RLHF with a fixed reward model. Overall, RLHF scales less efficiently than pretraining, with diminishing returns from additional computational resources. Based on these observations, we propose strategies to optimize RLHF performance within computational limits.
title Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.06000