Multiplayer Nash Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Fang, Huang, Xu, Xuan, Weihao, Zhang, Zhiwei, Xiao, Yijia, Wan, Guancheng, Li, Xiaomin, Hu, Bing, Xia, Peng, Leskovec, Jure, Choi, Yejin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914449845125120
author Wu, Fang
Huang, Xu
Xuan, Weihao
Zhang, Zhiwei
Xiao, Yijia
Wan, Guancheng
Li, Xiaomin
Hu, Bing
Xia, Peng
Leskovec, Jure
Choi, Yejin
author_facet Wu, Fang
Huang, Xu
Xuan, Weihao
Zhang, Zhiwei
Xiao, Yijia
Wan, Guancheng
Li, Xiaomin
Hu, Bing
Xia, Peng
Leskovec, Jure
Choi, Yejin
contents Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to Nash learning from human feedback (NLHF). While this perspective has inspired algorithms such as INPO, ONPO, and EGPO that offer strong theoretical and empirical guarantees, they remain fundamentally restricted to two-player interactions, introducing a single-opponent bias that fails to capture the full complexity of realistic preference structures. This work introduces Multiplayer Nash Preference Optimization (MNPO), a novel framework that generalizes NLHF to the multiplayer regime. It formulates alignment as an n-player game, where each policy competes against a population of opponents while being regularized toward a reference model. We demonstrate that MNPO inherits the equilibrium guarantees of two-player methods while enabling richer competitive dynamics and improved coverage of diverse preference structures. Comprehensive empirical evaluation shows that MNPO consistently outperforms existing NLHF baselines on instruction-following benchmarks, achieving superior alignment quality under heterogeneous annotator conditions and mixed-policy evaluation scenarios. Together, these results establish MNPO as a principled and scalable framework for aligning LLMs with complex, non-transitive human preferences. Code is available at: https://github.com/smiles724/MNPO
format Preprint
id arxiv_https___arxiv_org_abs_2509_23102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multiplayer Nash Preference Optimization
Wu, Fang
Huang, Xu
Xuan, Weihao
Zhang, Zhiwei
Xiao, Yijia
Wan, Guancheng
Li, Xiaomin
Hu, Bing
Xia, Peng
Leskovec, Jure
Choi, Yejin
Artificial Intelligence
Computation and Language
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to Nash learning from human feedback (NLHF). While this perspective has inspired algorithms such as INPO, ONPO, and EGPO that offer strong theoretical and empirical guarantees, they remain fundamentally restricted to two-player interactions, introducing a single-opponent bias that fails to capture the full complexity of realistic preference structures. This work introduces Multiplayer Nash Preference Optimization (MNPO), a novel framework that generalizes NLHF to the multiplayer regime. It formulates alignment as an n-player game, where each policy competes against a population of opponents while being regularized toward a reference model. We demonstrate that MNPO inherits the equilibrium guarantees of two-player methods while enabling richer competitive dynamics and improved coverage of diverse preference structures. Comprehensive empirical evaluation shows that MNPO consistently outperforms existing NLHF baselines on instruction-following benchmarks, achieving superior alignment quality under heterogeneous annotator conditions and mixed-policy evaluation scenarios. Together, these results establish MNPO as a principled and scalable framework for aligning LLMs with complex, non-transitive human preferences. Code is available at: https://github.com/smiles724/MNPO
title Multiplayer Nash Preference Optimization
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.23102