Learn Your Reference Model for Real Good Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gorbatovski, Alexey, Shaposhnikov, Boris, Malakhov, Alexey, Surnachev, Nikita, Aksenov, Yaroslav, Maksimov, Ian, Balagansky, Nikita, Gavrilov, Daniil
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909509820088320
author Gorbatovski, Alexey
Shaposhnikov, Boris
Malakhov, Alexey
Surnachev, Nikita
Aksenov, Yaroslav
Maksimov, Ian
Balagansky, Nikita
Gavrilov, Daniil
author_facet Gorbatovski, Alexey
Shaposhnikov, Boris
Malakhov, Alexey
Surnachev, Nikita
Aksenov, Yaroslav
Maksimov, Ian
Balagansky, Nikita
Gavrilov, Daniil
contents Despite the fact that offline methods for Large Language Models (LLMs) alignment do not require a direct reward model, they remain susceptible to overoptimization. This issue arises when the trained model deviates excessively from the reference policy, leading to a decrease in sample quality. We propose a new paradigm of offline alignment methods, called Trust Region (including variants TR-DPO, TR-IPO, TR-KTO), which dynamically updates the reference policy throughout the training process. Our results show that TR alignment methods effectively mitigate overoptimization, enabling models to maintain strong performance even when substantially deviating from the initial reference policy. We demonstrate the efficacy of these approaches not only through toy examples that exhibit reduced overoptimization, but also through direct, side-by-side comparisons in specific tasks such as helpful and harmless dialogue, as well as summarization, where they surpass conventional methods. Additionally, we report significant improvements in general-purpose assistant setups with the Llama3 model on the AlpacaEval 2 and Arena-Hard benchmarks, highlighting the advantages of Trust Region methods over classical approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09656
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learn Your Reference Model for Real Good Alignment
Gorbatovski, Alexey
Shaposhnikov, Boris
Malakhov, Alexey
Surnachev, Nikita
Aksenov, Yaroslav
Maksimov, Ian
Balagansky, Nikita
Gavrilov, Daniil
Machine Learning
Computation and Language
Despite the fact that offline methods for Large Language Models (LLMs) alignment do not require a direct reward model, they remain susceptible to overoptimization. This issue arises when the trained model deviates excessively from the reference policy, leading to a decrease in sample quality. We propose a new paradigm of offline alignment methods, called Trust Region (including variants TR-DPO, TR-IPO, TR-KTO), which dynamically updates the reference policy throughout the training process. Our results show that TR alignment methods effectively mitigate overoptimization, enabling models to maintain strong performance even when substantially deviating from the initial reference policy. We demonstrate the efficacy of these approaches not only through toy examples that exhibit reduced overoptimization, but also through direct, side-by-side comparisons in specific tasks such as helpful and harmless dialogue, as well as summarization, where they surpass conventional methods. Additionally, we report significant improvements in general-purpose assistant setups with the Llama3 model on the AlpacaEval 2 and Arena-Hard benchmarks, highlighting the advantages of Trust Region methods over classical approaches.
title Learn Your Reference Model for Real Good Alignment
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2404.09656