Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yuheng, Yu, Dian, Peng, Baolin, Song, Linfeng, Tian, Ye, Huo, Mingyue, Jiang, Nan, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
by: Yu, Dian, et al.
Published: (2024)
by: Yu, Dian, et al.
Published: (2024)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
by: Yu, Eason, et al.
Published: (2025)
by: Yu, Eason, et al.
Published: (2025)
Nash Equilibrium in Games on Graphs with Incomplete Preferences
by: Kulkarni, Abhishek N., et al.
Published: (2024)
by: Kulkarni, Abhishek N., et al.
Published: (2024)
Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
by: Liu, Kaizhao, et al.
Published: (2025)
by: Liu, Kaizhao, et al.
Published: (2025)
Teaching LLMs to Refine with Tools
by: Yu, Dian, et al.
Published: (2024)
by: Yu, Dian, et al.
Published: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
by: Jin, Lifeng, et al.
Published: (2024)
by: Jin, Lifeng, et al.
Published: (2024)
Nash equilibria of quasisupermodular games
by: Yu, Lu
Published: (2024)
by: Yu, Lu
Published: (2024)
Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games
by: Dong, Jing, et al.
Published: (2024)
by: Dong, Jing, et al.
Published: (2024)
Last-Iterate Convergence in Adaptive Regret Minimization for Approximate Extensive-Form Perfect Equilibrium
by: Ren, Hang, et al.
Published: (2025)
by: Ren, Hang, et al.
Published: (2025)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
LiteSearch: Efficacious Tree Search for LLM
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
by: Liu, Junyan, et al.
Published: (2026)
by: Liu, Junyan, et al.
Published: (2026)
On the Limitations and Possibilities of Nash Regret Minimization in Zero-Sum Matrix Games under Noisy Feedback
by: Maiti, Arnab, et al.
Published: (2023)
by: Maiti, Arnab, et al.
Published: (2023)
Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
by: Cai, Yang, et al.
Published: (2023)
by: Cai, Yang, et al.
Published: (2023)
Last-Iterate Convergence of No-Regret Learning for Equilibria in Bargaining Games
by: Kamp, Serafina, et al.
Published: (2025)
by: Kamp, Serafina, et al.
Published: (2025)
Preference-CFR$\:$ Beyond Nash Equilibrium for Better Game Strategies
by: Ju, Qi, et al.
Published: (2024)
by: Ju, Qi, et al.
Published: (2024)
Efficient Preference Elicitation in Iterative Combinatorial Auctions with Many Participants
by: Maruo, Ryota, et al.
Published: (2024)
by: Maruo, Ryota, et al.
Published: (2024)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Nash CoT: Multi-Path Inference with Preference Equilibrium
by: Zhang, Ziqi, et al.
Published: (2024)
by: Zhang, Ziqi, et al.
Published: (2024)
Efficient Last-Iterate Convergence in Regret Minimization via Adaptive Reward Transformation
by: Ren, Hang, et al.
Published: (2025)
by: Ren, Hang, et al.
Published: (2025)
Policy Abstraction and Nash Refinement in Tree-Exploiting PSRO
by: Konicki, Christine, et al.
Published: (2025)
by: Konicki, Christine, et al.
Published: (2025)
Compression-based Privacy Preservation for Distributed Nash Equilibrium Seeking in Aggregative Games
by: Huo, Wei, et al.
Published: (2024)
by: Huo, Wei, et al.
Published: (2024)
Communication-efficient and Differentially-private Distributed Nash Equilibrium Seeking with Linear Convergence
by: Chen, Xiaomeng, et al.
Published: (2024)
by: Chen, Xiaomeng, et al.
Published: (2024)
Evolution of Preferences in Multiple Populations
by: Tu, Yu-Sung, et al.
Published: (2018)
by: Tu, Yu-Sung, et al.
Published: (2018)
Convergence of Fast Policy Iteration in Markov Games and Robust MDPs
by: Badger, Keith, et al.
Published: (2025)
by: Badger, Keith, et al.
Published: (2025)
On the Complexity of Learning Nash Equilibria
by: Biggar, Oliver, et al.
Published: (2026)
by: Biggar, Oliver, et al.
Published: (2026)
Reinforcement Nash Equilibrium Solver
by: Wang, Xinrun, et al.
Published: (2024)
by: Wang, Xinrun, et al.
Published: (2024)
Satisfaction and Regret in Stackelberg Games
by: White, Langford, et al.
Published: (2024)
by: White, Langford, et al.
Published: (2024)
Nash Equilibrium and Belief Evolution in Differential Games
by: Zhou, Jiangjing, et al.
Published: (2025)
by: Zhou, Jiangjing, et al.
Published: (2025)
Beyond Pessimism: Offline Learning in KL-regularized Games
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
Nash Equilibria with Derangement Degree Probabilities
by: Orzech, Edan, et al.
Published: (2026)
by: Orzech, Edan, et al.
Published: (2026)
Randomness Requirements and Asymmetries in Nash Equilibria
by: Orzech, Edan, et al.
Published: (2023)
by: Orzech, Edan, et al.
Published: (2023)
Matching with Committee Preferences
by: Song, Haoyu, et al.
Published: (2026)
by: Song, Haoyu, et al.
Published: (2026)
Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making
by: Zheng, Kehan, et al.
Published: (2025)
by: Zheng, Kehan, et al.
Published: (2025)
On Optimal Tradeoffs between EFX and Nash Welfare
by: Feldman, Michal, et al.
Published: (2023)
by: Feldman, Michal, et al.
Published: (2023)
BAR Nash Equilibrium and Application to Blockchain Design
by: Reynouard, Maxime, et al.
Published: (2024)
by: Reynouard, Maxime, et al.
Published: (2024)
Nash Equilibria in Reverse Temporal Voronoi Games
by: Pawlowski, Simeon, et al.
Published: (2024)
by: Pawlowski, Simeon, et al.
Published: (2024)
On the Uniqueness of Nash Equilibria in Multiagent Matrix Games
by: Bailey, James P.
Published: (2024)
by: Bailey, James P.
Published: (2024)
Similar Items
-
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024) -
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
by: Yu, Dian, et al.
Published: (2024) -
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
by: Tian, Ye, et al.
Published: (2024) -
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
by: Yu, Eason, et al.
Published: (2025) -
Nash Equilibrium in Games on Graphs with Incomplete Preferences
by: Kulkarni, Abhishek N., et al.
Published: (2024)