Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Barla, Adam, Nevali, Emanuele, Viano, Luca, Cevher, Volkan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916015121629184
author Barla, Adam
Nevali, Emanuele
Viano, Luca
Cevher, Volkan
author_facet Barla, Adam
Nevali, Emanuele
Viano, Luca
Cevher, Volkan
contents We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigate the well-known over-optimization issue in preference learning without requiring the knowledge of the data-generating distribution or learning an explicit reward model. PEPO achieves pessimism via an ensemble of preference-optimized policies trained on disjoint data subsets and then aggregates them through a worst case construction that favors the agreement across models. In the tabular setting, PEPO achieves sample complexity guarantees depending only on a single-policy concentrability coefficient, thus avoiding the all-policy concentrability which affects the guarantees of algorithms prone to over-optimization, such as DPO. The theoretical findings are corroborated by a convincing practical performance, while retaining the simplicity and the practicality of DPO-style training.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06239
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
Barla, Adam
Nevali, Emanuele
Viano, Luca
Cevher, Volkan
Machine Learning
We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigate the well-known over-optimization issue in preference learning without requiring the knowledge of the data-generating distribution or learning an explicit reward model. PEPO achieves pessimism via an ensemble of preference-optimized policies trained on disjoint data subsets and then aggregates them through a worst case construction that favors the agreement across models. In the tabular setting, PEPO achieves sample complexity guarantees depending only on a single-policy concentrability coefficient, thus avoiding the all-policy concentrability which affects the guarantees of algorithms prone to over-optimization, such as DPO. The theoretical findings are corroborated by a convincing practical performance, while retaining the simplicity and the practicality of DPO-style training.
title Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
topic Machine Learning
url https://arxiv.org/abs/2602.06239