Jackpot! Alignment as a Maximal Lottery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maura-Rivero, Roberto-Rafael, Lanctot, Marc, Visin, Francesco, Larson, Kate
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916592456040448
author Maura-Rivero, Roberto-Rafael
Lanctot, Marc
Visin, Francesco
Larson, Kate
author_facet Maura-Rivero, Roberto-Rafael
Lanctot, Marc
Visin, Francesco
Larson, Kate
contents Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuitively desirable, such as respecting the preferences of the majority \cite{ge2024axioms}. To overcome these issues, we propose the use of a probabilistic Social Choice rule called \emph{maximal lotteries} as a replacement for RLHF. We show that a family of alignment techniques, namely Nash Learning from Human Feedback (NLHF) \cite{munos2023nash} and variants, approximate maximal lottery outcomes and thus inherit its beneficial properties. We confirm experimentally that our proposed methodology handles situations that arise when working with preferences more robustly than standard RLHF, including supporting the preferences of the majority, providing principled ways of handling non-transitivities in the preference data, and robustness to irrelevant alternatives. This results in systems that better incorporate human values and respect human intentions.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jackpot! Alignment as a Maximal Lottery
Maura-Rivero, Roberto-Rafael
Lanctot, Marc
Visin, Francesco
Larson, Kate
Artificial Intelligence
Machine Learning
Theoretical Economics
Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuitively desirable, such as respecting the preferences of the majority \cite{ge2024axioms}. To overcome these issues, we propose the use of a probabilistic Social Choice rule called \emph{maximal lotteries} as a replacement for RLHF. We show that a family of alignment techniques, namely Nash Learning from Human Feedback (NLHF) \cite{munos2023nash} and variants, approximate maximal lottery outcomes and thus inherit its beneficial properties. We confirm experimentally that our proposed methodology handles situations that arise when working with preferences more robustly than standard RLHF, including supporting the preferences of the majority, providing principled ways of handling non-transitivities in the preference data, and robustness to irrelevant alternatives. This results in systems that better incorporate human values and respect human intentions.
title Jackpot! Alignment as a Maximal Lottery
topic Artificial Intelligence
Machine Learning
Theoretical Economics
url https://arxiv.org/abs/2501.19266