Prior Constraints-based Reward Model Training for Aligning Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Hang, Wang, Chenglong, Hu, Yimin, Xiao, Tong, Zhang, Chunliang, Zhu, Jingbo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914951439843328
author Zhou, Hang
Wang, Chenglong
Hu, Yimin
Xiao, Tong
Zhang, Chunliang
Zhu, Jingbo
author_facet Zhou, Hang
Wang, Chenglong
Hu, Yimin
Xiao, Tong
Zhang, Chunliang
Zhu, Jingbo
contents Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training procedure suffers from an inherent problem: the uncontrolled scaling of reward scores during reinforcement learning due to the lack of constraints while training the reward model.This paper proposes a Prior Constraints-based Reward Model (namely PCRM) training method to mitigate this problem. PCRM incorporates prior constraints, specifically, length ratio and cosine similarity between outputs of each comparison pair, during reward model training to regulate optimization magnitude and control score margins. We comprehensively evaluate PCRM by examining its rank correlation with human preferences and its effectiveness in aligning LLMs via RL. Experimental results demonstrate that PCRM significantly improves alignment performance by effectively constraining reward score scaling. As another bonus, our method is easily integrated into arbitrary rank-based alignment methods, such as direct preference optimization, and can yield consistent improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00978
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prior Constraints-based Reward Model Training for Aligning Large Language Models
Zhou, Hang
Wang, Chenglong
Hu, Yimin
Xiao, Tong
Zhang, Chunliang
Zhu, Jingbo
Computation and Language
Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training procedure suffers from an inherent problem: the uncontrolled scaling of reward scores during reinforcement learning due to the lack of constraints while training the reward model.This paper proposes a Prior Constraints-based Reward Model (namely PCRM) training method to mitigate this problem. PCRM incorporates prior constraints, specifically, length ratio and cosine similarity between outputs of each comparison pair, during reward model training to regulate optimization magnitude and control score margins. We comprehensively evaluate PCRM by examining its rank correlation with human preferences and its effectiveness in aligning LLMs via RL. Experimental results demonstrate that PCRM significantly improves alignment performance by effectively constraining reward score scaling. As another bonus, our method is easily integrated into arbitrary rank-based alignment methods, such as direct preference optimization, and can yield consistent improvement.
title Prior Constraints-based Reward Model Training for Aligning Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2404.00978