Policy Optimization in Multi-Agent Settings under Partially Observable Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908482172616704 |
|---|---|
| author | Zhaikhan, Ainur Khammassi, Malek Sayed, Ali H. |
| author_facet | Zhaikhan, Ainur Khammassi, Malek Sayed, Ali H. |
| contents | This work leverages adaptive social learning to estimate partially observable global states in multi-agent reinforcement learning (MARL) problems. Unlike existing methods, the proposed approach enables the concurrent operation of social learning and reinforcement learning. Specifically, it alternates between a single step of social learning and a single step of MARL, eliminating the need for the time- and computation-intensive two-timescale learning frameworks. Theoretical guarantees are provided to support the effectiveness of the proposed method. Simulation results verify that the performance of the proposed methodology can approach that of reinforcement learning when the true state is known. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06061 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Policy Optimization in Multi-Agent Settings under Partially Observable Environments Zhaikhan, Ainur Khammassi, Malek Sayed, Ali H. Multiagent Systems This work leverages adaptive social learning to estimate partially observable global states in multi-agent reinforcement learning (MARL) problems. Unlike existing methods, the proposed approach enables the concurrent operation of social learning and reinforcement learning. Specifically, it alternates between a single step of social learning and a single step of MARL, eliminating the need for the time- and computation-intensive two-timescale learning frameworks. Theoretical guarantees are provided to support the effectiveness of the proposed method. Simulation results verify that the performance of the proposed methodology can approach that of reinforcement learning when the true state is known. |
| title | Policy Optimization in Multi-Agent Settings under Partially Observable Environments |
| topic | Multiagent Systems |
| url | https://arxiv.org/abs/2508.06061 |