A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Mingyang, Farina, Gabriele, Ozdaglar, Asuman
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916833410416640
author Liu, Mingyang
Farina, Gabriele
Ozdaglar, Asuman
author_facet Liu, Mingyang
Farina, Gabriele
Ozdaglar, Asuman
contents Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound avoidance of importance sampling corrections. In multi-agent imperfect-information settings (extensive-form games), however, it is still unknown whether the same desiderata can be guaranteed while retaining theoretical guarantees. Instead, sound methods for extensive-form games rely on approximating \emph{counterfactual} values (as opposed to Q values), which are incompatible with policy gradient methodologies. In this paper, we investigate whether policy gradient can be safely used in two-player zero-sum imperfect-information extensive-form games (EFGs). We establish positive results, showing for the first time that a policy gradient method leads to provable best-iterate convergence to a regularized Nash equilibrium in self-play.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
Liu, Mingyang
Farina, Gabriele
Ozdaglar, Asuman
Computer Science and Game Theory
Artificial Intelligence
Machine Learning
Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound avoidance of importance sampling corrections. In multi-agent imperfect-information settings (extensive-form games), however, it is still unknown whether the same desiderata can be guaranteed while retaining theoretical guarantees. Instead, sound methods for extensive-form games rely on approximating \emph{counterfactual} values (as opposed to Q values), which are incompatible with policy gradient methodologies. In this paper, we investigate whether policy gradient can be safely used in two-player zero-sum imperfect-information extensive-form games (EFGs). We establish positive results, showing for the first time that a policy gradient method leads to provable best-iterate convergence to a regularized Nash equilibrium in self-play.
title A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
topic Computer Science and Game Theory
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.00751