Scalable spectral representations for multi-agent reinforcement learning in network MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915024166977536 |
|---|---|
| author | Ren, Zhaolin Zhang, Runyu Dai, Bo Li, Na |
| author_facet | Ren, Zhaolin Zhang, Runyu Dai, Bo Li, Na |
| contents | Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global state-action space with the number of agents. In this work, utilizing the exponential decay property of network dynamics, we first derive scalable spectral local representations for network MDPs, which induces a network linear subspace for the local $Q$-function of each agent. Building on these local spectral representations, we design a scalable algorithmic framework for continuous state-action network MDPs, and provide end-to-end guarantees for the convergence of our algorithm. Empirically, we validate the effectiveness of our scalable representation-based approach on two benchmark problems, and demonstrate the advantages of our approach over generic function approximation approaches to representing the local $Q$-functions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_17221 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Scalable spectral representations for multi-agent reinforcement learning in network MDPs Ren, Zhaolin Zhang, Runyu Dai, Bo Li, Na Multiagent Systems Machine Learning Systems and Control Optimization and Control Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global state-action space with the number of agents. In this work, utilizing the exponential decay property of network dynamics, we first derive scalable spectral local representations for network MDPs, which induces a network linear subspace for the local $Q$-function of each agent. Building on these local spectral representations, we design a scalable algorithmic framework for continuous state-action network MDPs, and provide end-to-end guarantees for the convergence of our algorithm. Empirically, we validate the effectiveness of our scalable representation-based approach on two benchmark problems, and demonstrate the advantages of our approach over generic function approximation approaches to representing the local $Q$-functions. |
| title | Scalable spectral representations for multi-agent reinforcement learning in network MDPs |
| topic | Multiagent Systems Machine Learning Systems and Control Optimization and Control |
| url | https://arxiv.org/abs/2410.17221 |