NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918476147326976 |
|---|---|
| author | Yu, Eason Liu, Tzu Hao Canonne, Clément L. Wang, Yunke Xu, Chang Tran, Nguyen H. Albrecht, Stefano V. |
| author_facet | Yu, Eason Liu, Tzu Hao Canonne, Clément L. Wang, Yunke Xu, Chang Tran, Nguyen H. Albrecht, Stefano V. |
| contents | Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches either require full enumeration of the game tree or rely on non-policy-gradient inner solvers that underperform in practice, leaving a scalable policy-gradient-based solution open. In this paper, we propose a novel multi-round regularization procedure and show that it guarantees strictly monotonic reduction in Bregman divergence to Nash equilibria and eventual convergence to one in two-player zero-sum extensive-form games. Guided by this framework, we develop a practical algorithm, Nash Policy Gradient (NashPG), which places the regularization directly in the policy optimization objective and is implemented using standard policy gradient methods. Empirically, NashPG achieves comparable or lower exploitability than prior model-free methods on classic benchmark games and scales to large domains such as Battleship and No-Limit Texas Hold'em, where it attains higher average payoff in head-to-head play. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_18183 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Yu, Eason Liu, Tzu Hao Canonne, Clément L. Wang, Yunke Xu, Chang Tran, Nguyen H. Albrecht, Stefano V. Machine Learning Computer Science and Game Theory Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches either require full enumeration of the game tree or rely on non-policy-gradient inner solvers that underperform in practice, leaving a scalable policy-gradient-based solution open. In this paper, we propose a novel multi-round regularization procedure and show that it guarantees strictly monotonic reduction in Bregman divergence to Nash equilibria and eventual convergence to one in two-player zero-sum extensive-form games. Guided by this framework, we develop a practical algorithm, Nash Policy Gradient (NashPG), which places the regularization directly in the policy optimization objective and is implemented using standard policy gradient methods. Empirically, NashPG achieves comparable or lower exploitability than prior model-free methods on classic benchmark games and scales to large domains such as Battleship and No-Limit Texas Hold'em, where it attains higher average payoff in head-to-head play. |
| title | NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria |
| topic | Machine Learning Computer Science and Game Theory |
| url | https://arxiv.org/abs/2510.18183 |