Convergence of Decentralized Actor-Critic Algorithm in General-sum Markov Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maheshwari, Chinmay, Wu, Manxi, Sastry, Shankar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909646185299968
author Maheshwari, Chinmay
Wu, Manxi
Sastry, Shankar
author_facet Maheshwari, Chinmay
Wu, Manxi
Sastry, Shankar
contents Markov games provide a powerful framework for modeling strategic multi-agent interactions in dynamic environments. Traditionally, convergence properties of decentralized learning algorithms in these settings have been established only for special cases, such as Markov zero-sum and potential games, which do not fully capture real-world interactions. In this paper, we address this gap by studying the asymptotic properties of learning algorithms in general-sum Markov games. In particular, we focus on a decentralized algorithm where each agent adopts an actor-critic learning dynamic with asynchronous step sizes. This decentralized approach enables agents to operate independently, without requiring knowledge of others' strategies or payoffs. We introduce the concept of a Markov Near-Potential Function (MNPF) and demonstrate that it serves as an approximate Lyapunov function for the policy updates in the decentralized learning dynamics, which allows us to characterize the convergent set of strategies. We further strengthen our result under specific regularity conditions and with finite Nash equilibria.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Convergence of Decentralized Actor-Critic Algorithm in General-sum Markov Games
Maheshwari, Chinmay
Wu, Manxi
Sastry, Shankar
Multiagent Systems
Artificial Intelligence
Computer Science and Game Theory
Systems and Control
Optimization and Control
91A06, 91A10, 91A14, 91A15, 91A20, 91A40, 91A50, 93E03, 37N40
Markov games provide a powerful framework for modeling strategic multi-agent interactions in dynamic environments. Traditionally, convergence properties of decentralized learning algorithms in these settings have been established only for special cases, such as Markov zero-sum and potential games, which do not fully capture real-world interactions. In this paper, we address this gap by studying the asymptotic properties of learning algorithms in general-sum Markov games. In particular, we focus on a decentralized algorithm where each agent adopts an actor-critic learning dynamic with asynchronous step sizes. This decentralized approach enables agents to operate independently, without requiring knowledge of others' strategies or payoffs. We introduce the concept of a Markov Near-Potential Function (MNPF) and demonstrate that it serves as an approximate Lyapunov function for the policy updates in the decentralized learning dynamics, which allows us to characterize the convergent set of strategies. We further strengthen our result under specific regularity conditions and with finite Nash equilibria.
title Convergence of Decentralized Actor-Critic Algorithm in General-sum Markov Games
topic Multiagent Systems
Artificial Intelligence
Computer Science and Game Theory
Systems and Control
Optimization and Control
91A06, 91A10, 91A14, 91A15, 91A20, 91A40, 91A50, 93E03, 37N40
url https://arxiv.org/abs/2409.04613