Self-Evolved Reward Learning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918043640135680 |
|---|---|
| author | Huang, Chenghua Fan, Zhizhen Wang, Lu Yang, Fangkai Zhao, Pu Lin, Zeqi Lin, Qingwei Zhang, Dongmei Rajmohan, Saravan Zhang, Qi |
| author_facet | Huang, Chenghua Fan, Zhizhen Wang, Lu Yang, Fangkai Zhao, Pu Lin, Zeqi Lin, Qingwei Zhang, Dongmei Rajmohan, Saravan Zhang, Qi |
| contents | Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences, playing a pivotal role in the success of conversational models like GPT-4, ChatGPT, and Llama 2. A core challenge in employing RLHF lies in training a reliable reward model (RM), which relies on high-quality labels typically provided by human experts or advanced AI system. These methods can be costly and may introduce biases that affect the language model's responses. As language models improve, human input may become less effective in further enhancing their performance. In this paper, we propose Self-Evolved Reward Learning (SER), a novel approach where the RM generates additional training data to iteratively improve itself. We conducted extensive experiments on multiple datasets such as HH-RLHF and UltraFeedback, using models like Mistral and Llama 3, and compare SER against various baselines. Our results demonstrate that even with limited human-annotated data, learning from self-feedback can robustly enhance RM performance, thereby boosting the capabilities of large language models (LLMs). Resources of this paper can be found at https://aka.ms/ser |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_00418 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Self-Evolved Reward Learning for LLMs Huang, Chenghua Fan, Zhizhen Wang, Lu Yang, Fangkai Zhao, Pu Lin, Zeqi Lin, Qingwei Zhang, Dongmei Rajmohan, Saravan Zhang, Qi Computation and Language Artificial Intelligence Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences, playing a pivotal role in the success of conversational models like GPT-4, ChatGPT, and Llama 2. A core challenge in employing RLHF lies in training a reliable reward model (RM), which relies on high-quality labels typically provided by human experts or advanced AI system. These methods can be costly and may introduce biases that affect the language model's responses. As language models improve, human input may become less effective in further enhancing their performance. In this paper, we propose Self-Evolved Reward Learning (SER), a novel approach where the RM generates additional training data to iteratively improve itself. We conducted extensive experiments on multiple datasets such as HH-RLHF and UltraFeedback, using models like Mistral and Llama 3, and compare SER against various baselines. Our results demonstrate that even with limited human-annotated data, learning from self-feedback can robustly enhance RM performance, thereby boosting the capabilities of large language models (LLMs). Resources of this paper can be found at https://aka.ms/ser |
| title | Self-Evolved Reward Learning for LLMs |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2411.00418 |