Self-Evolved Reward Learning for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chenghua, Fan, Zhizhen, Wang, Lu, Yang, Fangkai, Zhao, Pu, Lin, Zeqi, Lin, Qingwei, Zhang, Dongmei, Rajmohan, Saravan, Zhang, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918043640135680
author Huang, Chenghua
Fan, Zhizhen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Lin, Zeqi
Lin, Qingwei
Zhang, Dongmei
Rajmohan, Saravan
Zhang, Qi
author_facet Huang, Chenghua
Fan, Zhizhen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Lin, Zeqi
Lin, Qingwei
Zhang, Dongmei
Rajmohan, Saravan
Zhang, Qi
contents Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences, playing a pivotal role in the success of conversational models like GPT-4, ChatGPT, and Llama 2. A core challenge in employing RLHF lies in training a reliable reward model (RM), which relies on high-quality labels typically provided by human experts or advanced AI system. These methods can be costly and may introduce biases that affect the language model's responses. As language models improve, human input may become less effective in further enhancing their performance. In this paper, we propose Self-Evolved Reward Learning (SER), a novel approach where the RM generates additional training data to iteratively improve itself. We conducted extensive experiments on multiple datasets such as HH-RLHF and UltraFeedback, using models like Mistral and Llama 3, and compare SER against various baselines. Our results demonstrate that even with limited human-annotated data, learning from self-feedback can robustly enhance RM performance, thereby boosting the capabilities of large language models (LLMs). Resources of this paper can be found at https://aka.ms/ser
format Preprint
id arxiv_https___arxiv_org_abs_2411_00418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Evolved Reward Learning for LLMs
Huang, Chenghua
Fan, Zhizhen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Lin, Zeqi
Lin, Qingwei
Zhang, Dongmei
Rajmohan, Saravan
Zhang, Qi
Computation and Language
Artificial Intelligence
Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences, playing a pivotal role in the success of conversational models like GPT-4, ChatGPT, and Llama 2. A core challenge in employing RLHF lies in training a reliable reward model (RM), which relies on high-quality labels typically provided by human experts or advanced AI system. These methods can be costly and may introduce biases that affect the language model's responses. As language models improve, human input may become less effective in further enhancing their performance. In this paper, we propose Self-Evolved Reward Learning (SER), a novel approach where the RM generates additional training data to iteratively improve itself. We conducted extensive experiments on multiple datasets such as HH-RLHF and UltraFeedback, using models like Mistral and Llama 3, and compare SER against various baselines. Our results demonstrate that even with limited human-annotated data, learning from self-feedback can robustly enhance RM performance, thereby boosting the capabilities of large language models (LLMs). Resources of this paper can be found at https://aka.ms/ser
title Self-Evolved Reward Learning for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.00418