Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Donghwan, Lim, Han-Dong, Kim, Do Wan
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:https://arxiv.org/abs/2307.16706
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910483960823808
author Lee, Donghwan
Lim, Han-Dong
Kim, Do Wan
author_facet Lee, Donghwan
Lim, Han-Dong
Kim, Do Wan
contents The main goal of this paper is to investigate continuous-time distributed dynamic programming (DP) algorithms for networked multi-agent Markov decision problems (MAMDPs). In our study, we adopt a distributed multi-agent framework where individual agents have access only to their own rewards, lacking insights into the rewards of other agents. Moreover, each agent has the ability to share its parameters with neighboring agents through a communication network, represented by a graph. We first introduce a novel distributed DP, inspired by the distributed optimization method of Wang and Elia. Next, a new distributed DP is introduced through a decoupling process. The convergence of the DP algorithms is proved through systems and control perspectives. The study in this paper sets the stage for new distributed temporal different learning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2307_16706
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes
Lee, Donghwan
Lim, Han-Dong
Kim, Do Wan
Systems and Control
Artificial Intelligence
The main goal of this paper is to investigate continuous-time distributed dynamic programming (DP) algorithms for networked multi-agent Markov decision problems (MAMDPs). In our study, we adopt a distributed multi-agent framework where individual agents have access only to their own rewards, lacking insights into the rewards of other agents. Moreover, each agent has the ability to share its parameters with neighboring agents through a communication network, represented by a graph. We first introduce a novel distributed DP, inspired by the distributed optimization method of Wang and Elia. Next, a new distributed DP is introduced through a decoupling process. The convergence of the DP algorithms is proved through systems and control perspectives. The study in this paper sets the stage for new distributed temporal different learning algorithms.
title Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes
topic Systems and Control
Artificial Intelligence
url https://arxiv.org/abs/2307.16706