Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909119912345600 |
|---|---|
| author | Mao, Xin Li, Feng-Lin Xu, Huimin Zhang, Wei Luu, Anh Tuan |
| author_facet | Mao, Xin Li, Feng-Lin Xu, Huimin Zhang, Wei Luu, Anh Tuan |
| contents | While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of order-based calibration methods as viable alternatives. This paper delves further into current order-based methods, examining their inefficiencies in utilizing reward values and addressing misalignment issues. Building upon these findings, we propose a novel \textbf{V}alue-based \textbf{C}ali\textbf{B}ration (VCB) method to better align LLMs with human preferences. Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and stability in diverse settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_16030 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration Mao, Xin Li, Feng-Lin Xu, Huimin Zhang, Wei Luu, Anh Tuan Computation and Language Artificial Intelligence While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of order-based calibration methods as viable alternatives. This paper delves further into current order-based methods, examining their inefficiencies in utilizing reward values and addressing misalignment issues. Building upon these findings, we propose a novel \textbf{V}alue-based \textbf{C}ali\textbf{B}ration (VCB) method to better align LLMs with human preferences. Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and stability in diverse settings. |
| title | Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2402.16030 |