The Sample-Communication Complexity Trade-off in Federated Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salgia, Sudeep, Chi, Yuejie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567211520000
author Salgia, Sudeep
Chi, Yuejie
author_facet Salgia, Sudeep
Chi, Yuejie
contents We consider the problem of federated Q-learning, where $M$ agents aim to collaboratively learn the optimal Q-function of an unknown infinite-horizon Markov decision process with finite state and action spaces. We investigate the trade-off between sample and communication complexities for the widely used class of intermittent communication algorithms. We first establish the converse result, where it is shown that a federated Q-learning algorithm that offers any speedup with respect to the number of agents in the per-agent sample complexity needs to incur a communication cost of at least an order of $\frac{1}{1-γ}$ up to logarithmic factors, where $γ$ is the discount factor. We also propose a new algorithm, called Fed-DVR-Q, which is the first federated Q-learning algorithm to simultaneously achieve order-optimal sample and communication complexities. Thus, together these results provide a complete characterization of the sample-communication complexity trade-off in federated Q-learning.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16981
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Sample-Communication Complexity Trade-off in Federated Q-Learning
Salgia, Sudeep
Chi, Yuejie
Machine Learning
Optimization and Control
We consider the problem of federated Q-learning, where $M$ agents aim to collaboratively learn the optimal Q-function of an unknown infinite-horizon Markov decision process with finite state and action spaces. We investigate the trade-off between sample and communication complexities for the widely used class of intermittent communication algorithms. We first establish the converse result, where it is shown that a federated Q-learning algorithm that offers any speedup with respect to the number of agents in the per-agent sample complexity needs to incur a communication cost of at least an order of $\frac{1}{1-γ}$ up to logarithmic factors, where $γ$ is the discount factor. We also propose a new algorithm, called Fed-DVR-Q, which is the first federated Q-learning algorithm to simultaneously achieve order-optimal sample and communication complexities. Thus, together these results provide a complete characterization of the sample-communication complexity trade-off in federated Q-learning.
title The Sample-Communication Complexity Trade-off in Federated Q-Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2408.16981