Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yongacoglu, Bora, Arslan, Gürdal, Yüksel, Serdar
Format: Preprint
Publié: 2021
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911814864863232
author Yongacoglu, Bora
Arslan, Gürdal
Yüksel, Serdar
author_facet Yongacoglu, Bora
Arslan, Gürdal
Yüksel, Serdar
contents In multi-agent reinforcement learning (MARL), independent learners are those that do not observe the actions of other agents in the system. Due to the decentralization of information, it is challenging to design independent learners that drive play to equilibrium. This paper investigates the feasibility of using satisficing dynamics to guide independent learners to approximate equilibrium in stochastic games. For $ε\geq 0$, an $ε$-satisficing policy update rule is any rule that instructs the agent to not change its policy when it is $ε$-best-responding to the policies of the remaining players; $ε$-satisficing paths are defined to be sequences of joint policies obtained when each agent uses some $ε$-satisficing policy update rule to select its next policy. We establish structural results on the existence of $ε$-satisficing paths into $ε$-equilibrium in both symmetric $N$-player games and general stochastic games with two players. We then present an independent learning algorithm for $N$-player symmetric games and give high probability guarantees of convergence to $ε$-equilibrium under self-play. This guarantee is made using symmetry alone, leveraging the previously unexploited structure of $ε$-satisficing paths.
format Preprint
id arxiv_https___arxiv_org_abs_2110_04638
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games
Yongacoglu, Bora
Arslan, Gürdal
Yüksel, Serdar
Computer Science and Game Theory
Machine Learning
Optimization and Control
In multi-agent reinforcement learning (MARL), independent learners are those that do not observe the actions of other agents in the system. Due to the decentralization of information, it is challenging to design independent learners that drive play to equilibrium. This paper investigates the feasibility of using satisficing dynamics to guide independent learners to approximate equilibrium in stochastic games. For $ε\geq 0$, an $ε$-satisficing policy update rule is any rule that instructs the agent to not change its policy when it is $ε$-best-responding to the policies of the remaining players; $ε$-satisficing paths are defined to be sequences of joint policies obtained when each agent uses some $ε$-satisficing policy update rule to select its next policy. We establish structural results on the existence of $ε$-satisficing paths into $ε$-equilibrium in both symmetric $N$-player games and general stochastic games with two players. We then present an independent learning algorithm for $N$-player symmetric games and give high probability guarantees of convergence to $ε$-equilibrium under self-play. This guarantee is made using symmetry alone, leveraging the previously unexploited structure of $ε$-satisficing paths.
title Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games
topic Computer Science and Game Theory
Machine Learning
Optimization and Control
url https://arxiv.org/abs/2110.04638