Geometric-Averaged Preference Optimization for Soft Preference Labels

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Furuta, Hiroki, Lee, Kuang-Huei, Gu, Shixiang Shane, Matsuo, Yutaka, Faust, Aleksandra, Zen, Heiga, Gur, Izzeddin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916544295993344
author Furuta, Hiroki
Lee, Kuang-Huei
Gu, Shixiang Shane
Matsuo, Yutaka
Faust, Aleksandra
Zen, Heiga
Gur, Izzeddin
author_facet Furuta, Hiroki
Lee, Kuang-Huei
Gu, Shixiang Shane
Matsuo, Yutaka
Faust, Aleksandra
Zen, Heiga
Gur, Izzeddin
contents Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and therefore should be represented distributionally. In this work, we introduce the distributional soft preference labels and improve Direct Preference Optimization (DPO) with a weighted geometric average of the LLM output likelihood in the loss function. This approach adjusts the scale of learning loss based on the soft labels such that the loss would approach zero when the responses are closer to equally preferred. This simple modification can be easily applied to any DPO-based methods and mitigate over-optimization and objective mismatch, which prior works suffer from. Our experiments simulate the soft preference labels with AI feedback from LLMs and demonstrate that geometric averaging consistently improves performance on standard benchmarks for alignment research. In particular, we observe more preferable responses than binary labels and significant improvements where modestly-confident labels are in the majority.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06691
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Geometric-Averaged Preference Optimization for Soft Preference Labels
Furuta, Hiroki
Lee, Kuang-Huei
Gu, Shixiang Shane
Matsuo, Yutaka
Faust, Aleksandra
Zen, Heiga
Gur, Izzeddin
Machine Learning
Artificial Intelligence
Computation and Language
Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and therefore should be represented distributionally. In this work, we introduce the distributional soft preference labels and improve Direct Preference Optimization (DPO) with a weighted geometric average of the LLM output likelihood in the loss function. This approach adjusts the scale of learning loss based on the soft labels such that the loss would approach zero when the responses are closer to equally preferred. This simple modification can be easily applied to any DPO-based methods and mitigate over-optimization and objective mismatch, which prior works suffer from. Our experiments simulate the soft preference labels with AI feedback from LLMs and demonstrate that geometric averaging consistently improves performance on standard benchmarks for alignment research. In particular, we observe more preferable responses than binary labels and significant improvements where modestly-confident labels are in the majority.
title Geometric-Averaged Preference Optimization for Soft Preference Labels
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.06691