The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Shengyi, Noukhovitch, Michael, Hosseini, Arian, Rasul, Kashif, Wang, Weixun, Tunstall, Lewis
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911812865228800
author Huang, Shengyi
Noukhovitch, Michael
Hosseini, Arian
Rasul, Kashif
Wang, Weixun
Tunstall, Lewis
author_facet Huang, Shengyi
Noukhovitch, Michael
Hosseini, Arian
Rasul, Kashif
Wang, Weixun
Tunstall, Lewis
contents This work is the first to openly reproduce the Reinforcement Learning from Human Feedback (RLHF) scaling behaviors reported in OpenAI's seminal TL;DR summarization work. We create an RLHF pipeline from scratch, enumerate over 20 key implementation details, and share key insights during the reproduction. Our RLHF-trained Pythia models demonstrate significant gains in response quality that scale with model size, with our 2.8B, 6.9B models outperforming OpenAI's released 1.3B checkpoint. We publicly release the trained model checkpoints and code to facilitate further research and accelerate progress in the field (\url{https://github.com/vwxyzjn/summarize_from_feedback_details}).
format Preprint
id arxiv_https___arxiv_org_abs_2403_17031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Huang, Shengyi
Noukhovitch, Michael
Hosseini, Arian
Rasul, Kashif
Wang, Weixun
Tunstall, Lewis
Machine Learning
This work is the first to openly reproduce the Reinforcement Learning from Human Feedback (RLHF) scaling behaviors reported in OpenAI's seminal TL;DR summarization work. We create an RLHF pipeline from scratch, enumerate over 20 key implementation details, and share key insights during the reproduction. Our RLHF-trained Pythia models demonstrate significant gains in response quality that scale with model size, with our 2.8B, 6.9B models outperforming OpenAI's released 1.3B checkpoint. We publicly release the trained model checkpoints and code to facilitate further research and accelerate progress in the field (\url{https://github.com/vwxyzjn/summarize_from_feedback_details}).
title The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
topic Machine Learning
url https://arxiv.org/abs/2403.17031