A finite time analysis of distributed Q-learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lim, Han-Dong, Lee, Donghwan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911080206303232
author Lim, Han-Dong
Lee, Donghwan
author_facet Lim, Han-Dong
Lee, Donghwan
contents Multi-agent reinforcement learning (MARL) has witnessed a remarkable surge in interest, fueled by the empirical success achieved in applications of single-agent reinforcement learning (RL). In this study, we consider a distributed Q-learning scenario, wherein a number of agents cooperatively solve a sequential decision making problem without access to the central reward function which is an average of the local rewards. In particular, we study finite-time analysis of a distributed Q-learning algorithm, and provide a new sample complexity result of $\tilde{\mathcal{O}}\left( \min\left\{\frac{1}{ε^2}\frac{t_{\text{mix}}}{(1-γ)^6 d_{\min}^4 } ,\frac{1}ε\frac{\sqrt{|\gS||\gA|}}{(1-σ_2(\boldsymbol{W}))(1-γ)^4 d_{\min}^3} \right\}\right)$ under tabular lookup
format Preprint
id arxiv_https___arxiv_org_abs_2405_14078
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A finite time analysis of distributed Q-learning
Lim, Han-Dong
Lee, Donghwan
Artificial Intelligence
Machine Learning
Multiagent Systems
Multi-agent reinforcement learning (MARL) has witnessed a remarkable surge in interest, fueled by the empirical success achieved in applications of single-agent reinforcement learning (RL). In this study, we consider a distributed Q-learning scenario, wherein a number of agents cooperatively solve a sequential decision making problem without access to the central reward function which is an average of the local rewards. In particular, we study finite-time analysis of a distributed Q-learning algorithm, and provide a new sample complexity result of $\tilde{\mathcal{O}}\left( \min\left\{\frac{1}{ε^2}\frac{t_{\text{mix}}}{(1-γ)^6 d_{\min}^4 } ,\frac{1}ε\frac{\sqrt{|\gS||\gA|}}{(1-σ_2(\boldsymbol{W}))(1-γ)^4 d_{\min}^3} \right\}\right)$ under tabular lookup
title A finite time analysis of distributed Q-learning
topic Artificial Intelligence
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2405.14078