Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Jiesong, Huang, Yucong, Ma, Chengdong, Wang, Mingzhi, Wen, Ying, Hu, Long, Hao, Yixue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918270024548352
author Lian, Jiesong
Huang, Yucong
Ma, Chengdong
Wang, Mingzhi
Wen, Ying
Hu, Long
Hao, Yixue
author_facet Lian, Jiesong
Huang, Yucong
Ma, Chengdong
Wang, Mingzhi
Wen, Ying
Hu, Long
Hao, Yixue
contents For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy in Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.
format Preprint
id arxiv_https___arxiv_org_abs_2405_21027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
Lian, Jiesong
Huang, Yucong
Ma, Chengdong
Wang, Mingzhi
Wen, Ying
Hu, Long
Hao, Yixue
Computer Science and Game Theory
Artificial Intelligence
Machine Learning
Multiagent Systems
For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy in Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.
title Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
topic Computer Science and Game Theory
Artificial Intelligence
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2405.21027