Reevaluating Policy Gradient Methods for Imperfect-Information Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rudolph, Max, Lichtle, Nathan, Mohammadpour, Sobhan, Bayen, Alexandre, Kolter, J. Zico, Zhang, Amy, Farina, Gabriele, Vinitsky, Eugene, Sokota, Samuel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916051310084096
author Rudolph, Max
Lichtle, Nathan
Mohammadpour, Sobhan
Bayen, Alexandre
Kolter, J. Zico
Zhang, Amy
Farina, Gabriele
Vinitsky, Eugene
Sokota, Samuel
author_facet Rudolph, Max
Lichtle, Nathan
Mohammadpour, Sobhan
Bayen, Alexandre
Kolter, J. Zico
Zhang, Amy
Farina, Gabriele
Vinitsky, Eugene
Sokota, Samuel
contents In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for five large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algorithms for imperfect-information games. Over 7000 training runs, we find that FP-, DO-, and CFR-based approaches fail to outperform generic policy gradient methods. Code is available at https://github.com/nathanlct/IIG-RL-Benchmark and https://github.com/gabrfarina/exp-a-spiel .
format Preprint
id arxiv_https___arxiv_org_abs_2502_08938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reevaluating Policy Gradient Methods for Imperfect-Information Games
Rudolph, Max
Lichtle, Nathan
Mohammadpour, Sobhan
Bayen, Alexandre
Kolter, J. Zico
Zhang, Amy
Farina, Gabriele
Vinitsky, Eugene
Sokota, Samuel
Machine Learning
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for five large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algorithms for imperfect-information games. Over 7000 training runs, we find that FP-, DO-, and CFR-based approaches fail to outperform generic policy gradient methods. Code is available at https://github.com/nathanlct/IIG-RL-Benchmark and https://github.com/gabrfarina/exp-a-spiel .
title Reevaluating Policy Gradient Methods for Imperfect-Information Games
topic Machine Learning
url https://arxiv.org/abs/2502.08938