NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Eason, Liu, Tzu Hao, Canonne, Clément L., Wang, Yunke, Xu, Chang, Tran, Nguyen H., Albrecht, Stefano V.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918476147326976
author Yu, Eason
Liu, Tzu Hao
Canonne, Clément L.
Wang, Yunke
Xu, Chang
Tran, Nguyen H.
Albrecht, Stefano V.
author_facet Yu, Eason
Liu, Tzu Hao
Canonne, Clément L.
Wang, Yunke
Xu, Chang
Tran, Nguyen H.
Albrecht, Stefano V.
contents Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches either require full enumeration of the game tree or rely on non-policy-gradient inner solvers that underperform in practice, leaving a scalable policy-gradient-based solution open. In this paper, we propose a novel multi-round regularization procedure and show that it guarantees strictly monotonic reduction in Bregman divergence to Nash equilibria and eventual convergence to one in two-player zero-sum extensive-form games. Guided by this framework, we develop a practical algorithm, Nash Policy Gradient (NashPG), which places the regularization directly in the policy optimization objective and is implemented using standard policy gradient methods. Empirically, NashPG achieves comparable or lower exploitability than prior model-free methods on classic benchmark games and scales to large domains such as Battleship and No-Limit Texas Hold'em, where it attains higher average payoff in head-to-head play.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
Yu, Eason
Liu, Tzu Hao
Canonne, Clément L.
Wang, Yunke
Xu, Chang
Tran, Nguyen H.
Albrecht, Stefano V.
Machine Learning
Computer Science and Game Theory
Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches either require full enumeration of the game tree or rely on non-policy-gradient inner solvers that underperform in practice, leaving a scalable policy-gradient-based solution open. In this paper, we propose a novel multi-round regularization procedure and show that it guarantees strictly monotonic reduction in Bregman divergence to Nash equilibria and eventual convergence to one in two-player zero-sum extensive-form games. Guided by this framework, we develop a practical algorithm, Nash Policy Gradient (NashPG), which places the regularization directly in the policy optimization objective and is implemented using standard policy gradient methods. Empirically, NashPG achieves comparable or lower exploitability than prior model-free methods on classic benchmark games and scales to large domains such as Battleship and No-Limit Texas Hold'em, where it attains higher average payoff in head-to-head play.
title NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
topic Machine Learning
Computer Science and Game Theory
url https://arxiv.org/abs/2510.18183