One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913372758343680 |
|---|---|
| author | Tian, Haoxing Paschalidis, Ioannis Ch. Olshevsky, Alex |
| author_facet | Tian, Haoxing Paschalidis, Ioannis Ch. Olshevsky, Alex |
| contents | We consider a distributed setup for reinforcement learning, where each agent has a copy of the same Markov Decision Process but transitions are sampled from the corresponding Markov chain independently by each agent. We show that in this setting, we can achieve a linear speedup for TD($λ$), a family of popular methods for policy evaluation, in the sense that $N$ agents can evaluate a policy $N$ times faster provided the target accuracy is small enough. Notably, this speedup is achieved by ``one shot averaging,'' a procedure where the agents run TD($λ$) with Markov sampling independently and only average their results after the final step. This significantly reduces the amount of communication required to achieve a linear speedup relative to previous work. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_08896 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling Tian, Haoxing Paschalidis, Ioannis Ch. Olshevsky, Alex Machine Learning Distributed, Parallel, and Cluster Computing We consider a distributed setup for reinforcement learning, where each agent has a copy of the same Markov Decision Process but transitions are sampled from the corresponding Markov chain independently by each agent. We show that in this setting, we can achieve a linear speedup for TD($λ$), a family of popular methods for policy evaluation, in the sense that $N$ agents can evaluate a policy $N$ times faster provided the target accuracy is small enough. Notably, this speedup is achieved by ``one shot averaging,'' a procedure where the agents run TD($λ$) with Markov sampling independently and only average their results after the final step. This significantly reduces the amount of communication required to achieve a linear speedup relative to previous work. |
| title | One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2403.08896 |