On Monotonicity in AI Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bareilles, Gilles, Fageot, Julien, Hoang, Lê-Nguyên, Blanchard, Peva, Bouaziz, Wassim, Rouault, Sébastien, El-Mhamdi, El-Mahdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918059839586304
author Bareilles, Gilles
Fageot, Julien
Hoang, Lê-Nguyên
Blanchard, Peva
Bouaziz, Wassim
Rouault, Sébastien
El-Mhamdi, El-Mahdi
author_facet Bareilles, Gilles
Fageot, Julien
Hoang, Lê-Nguyên
Blanchard, Peva
Bouaziz, Wassim
Rouault, Sébastien
El-Mhamdi, El-Mahdi
contents Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Monotonicity in AI Alignment
Bareilles, Gilles
Fageot, Julien
Hoang, Lê-Nguyên
Blanchard, Peva
Bouaziz, Wassim
Rouault, Sébastien
El-Mhamdi, El-Mahdi
Statistics Theory
Machine Learning
Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms.
title On Monotonicity in AI Alignment
topic Statistics Theory
Machine Learning
url https://arxiv.org/abs/2506.08998