Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mao, Xin, Li, Feng-Lin, Xu, Huimin, Zhang, Wei, Luu, Anh Tuan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909119912345600
author Mao, Xin
Li, Feng-Lin
Xu, Huimin
Zhang, Wei
Luu, Anh Tuan
author_facet Mao, Xin
Li, Feng-Lin
Xu, Huimin
Zhang, Wei
Luu, Anh Tuan
contents While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of order-based calibration methods as viable alternatives. This paper delves further into current order-based methods, examining their inefficiencies in utilizing reward values and addressing misalignment issues. Building upon these findings, we propose a novel \textbf{V}alue-based \textbf{C}ali\textbf{B}ration (VCB) method to better align LLMs with human preferences. Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and stability in diverse settings.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16030
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
Mao, Xin
Li, Feng-Lin
Xu, Huimin
Zhang, Wei
Luu, Anh Tuan
Computation and Language
Artificial Intelligence
While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of order-based calibration methods as viable alternatives. This paper delves further into current order-based methods, examining their inefficiencies in utilizing reward values and addressing misalignment issues. Building upon these findings, we propose a novel \textbf{V}alue-based \textbf{C}ali\textbf{B}ration (VCB) method to better align LLMs with human preferences. Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and stability in diverse settings.
title Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.16030