Independent and Decentralized Learning in Markov Potential Games

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Maheshwari, Chinmay, Wu, Manxi, Pai, Druv, Sastry, Shankar
Formato: Preprint
Publicado: 2022
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912301926318080
author Maheshwari, Chinmay
Wu, Manxi
Pai, Druv
Sastry, Shankar
author_facet Maheshwari, Chinmay
Wu, Manxi
Pai, Druv
Sastry, Shankar
contents We study a multi-agent reinforcement learning dynamics, and analyze its asymptotic behavior in infinite-horizon discounted Markov potential games. We focus on the independent and decentralized setting, where players do not know the game parameters, and cannot communicate or coordinate. In each stage, players update their estimate of Q-function that evaluates their total contingent payoff based on the realized one-stage reward in an asynchronous manner. Then, players independently update their policies by incorporating an optimal one-stage deviation strategy based on the estimated Q-function. Inspired by the actor-critic algorithm in single-agent reinforcement learning, a key feature of our learning dynamics is that agents update their Q-function estimates at a faster timescale than the policies. Leveraging tools from two-timescale asynchronous stochastic approximation theory, we characterize the convergent set of learning dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2205_14590
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Independent and Decentralized Learning in Markov Potential Games
Maheshwari, Chinmay
Wu, Manxi
Pai, Druv
Sastry, Shankar
Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
Systems and Control
91A06, 91A10, 91A14, 91A25, 91A26, 91A50,
We study a multi-agent reinforcement learning dynamics, and analyze its asymptotic behavior in infinite-horizon discounted Markov potential games. We focus on the independent and decentralized setting, where players do not know the game parameters, and cannot communicate or coordinate. In each stage, players update their estimate of Q-function that evaluates their total contingent payoff based on the realized one-stage reward in an asynchronous manner. Then, players independently update their policies by incorporating an optimal one-stage deviation strategy based on the estimated Q-function. Inspired by the actor-critic algorithm in single-agent reinforcement learning, a key feature of our learning dynamics is that agents update their Q-function estimates at a faster timescale than the policies. Leveraging tools from two-timescale asynchronous stochastic approximation theory, we characterize the convergent set of learning dynamics.
title Independent and Decentralized Learning in Markov Potential Games
topic Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
Systems and Control
91A06, 91A10, 91A14, 91A25, 91A26, 91A50,
url https://arxiv.org/abs/2205.14590