On Monotonicity in AI Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918059839586304 |
|---|---|
| author | Bareilles, Gilles Fageot, Julien Hoang, Lê-Nguyên Blanchard, Peva Bouaziz, Wassim Rouault, Sébastien El-Mhamdi, El-Mahdi |
| author_facet | Bareilles, Gilles Fageot, Julien Hoang, Lê-Nguyên Blanchard, Peva Bouaziz, Wassim Rouault, Sébastien El-Mhamdi, El-Mahdi |
| contents | Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08998 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | On Monotonicity in AI Alignment Bareilles, Gilles Fageot, Julien Hoang, Lê-Nguyên Blanchard, Peva Bouaziz, Wassim Rouault, Sébastien El-Mhamdi, El-Mahdi Statistics Theory Machine Learning Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms. |
| title | On Monotonicity in AI Alignment |
| topic | Statistics Theory Machine Learning |
| url | https://arxiv.org/abs/2506.08998 |