A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916833410416640 |
|---|---|
| author | Liu, Mingyang Farina, Gabriele Ozdaglar, Asuman |
| author_facet | Liu, Mingyang Farina, Gabriele Ozdaglar, Asuman |
| contents | Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound avoidance of importance sampling corrections. In multi-agent imperfect-information settings (extensive-form games), however, it is still unknown whether the same desiderata can be guaranteed while retaining theoretical guarantees. Instead, sound methods for extensive-form games rely on approximating \emph{counterfactual} values (as opposed to Q values), which are incompatible with policy gradient methodologies. In this paper, we investigate whether policy gradient can be safely used in two-player zero-sum imperfect-information extensive-form games (EFGs). We establish positive results, showing for the first time that a policy gradient method leads to provable best-iterate convergence to a regularized Nash equilibrium in self-play. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_00751 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence Liu, Mingyang Farina, Gabriele Ozdaglar, Asuman Computer Science and Game Theory Artificial Intelligence Machine Learning Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound avoidance of importance sampling corrections. In multi-agent imperfect-information settings (extensive-form games), however, it is still unknown whether the same desiderata can be guaranteed while retaining theoretical guarantees. Instead, sound methods for extensive-form games rely on approximating \emph{counterfactual} values (as opposed to Q values), which are incompatible with policy gradient methodologies. In this paper, we investigate whether policy gradient can be safely used in two-player zero-sum imperfect-information extensive-form games (EFGs). We establish positive results, showing for the first time that a policy gradient method leads to provable best-iterate convergence to a regularized Nash equilibrium in self-play. |
| title | A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence |
| topic | Computer Science and Game Theory Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2408.00751 |