Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Gallego, Víctor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909105029906432
author Gallego, Víctor
author_facet Gallego, Víctor
contents In this paper, we introduce \emph{refined Direct Preference Optimization} (rDPO), a method for improving the behavioral alignment of Large Language Models (LLMs) without the need for human-annotated data. The method involves creating synthetic data using self-critique prompting by a teacher LLM and then utilising a generalized DPO loss function to distil to a student LLM. The loss function incorporates an additional external reward model to improve the quality of synthetic data, making rDPO robust to potential noise in the synthetic dataset. rDPO is shown to be effective in a diverse set of behavioural alignment tasks, such as improved safety, robustness against role-playing, and reduced sycophancy. Code to be released at https://github.com/vicgalle/refined-dpo.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08005
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
Gallego, Víctor
Computation and Language
Machine Learning
In this paper, we introduce \emph{refined Direct Preference Optimization} (rDPO), a method for improving the behavioral alignment of Large Language Models (LLMs) without the need for human-annotated data. The method involves creating synthetic data using self-critique prompting by a teacher LLM and then utilising a generalized DPO loss function to distil to a student LLM. The loss function incorporates an additional external reward model to improve the quality of synthetic data, making rDPO robust to potential noise in the synthetic dataset. rDPO is shown to be effective in a diverse set of behavioural alignment tasks, such as improved safety, robustness against role-playing, and reduced sycophancy. Code to be released at https://github.com/vicgalle/refined-dpo.
title Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.08005