Policy Optimization in Multi-Agent Settings under Partially Observable Environments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhaikhan, Ainur, Khammassi, Malek, Sayed, Ali H.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908482172616704
author Zhaikhan, Ainur
Khammassi, Malek
Sayed, Ali H.
author_facet Zhaikhan, Ainur
Khammassi, Malek
Sayed, Ali H.
contents This work leverages adaptive social learning to estimate partially observable global states in multi-agent reinforcement learning (MARL) problems. Unlike existing methods, the proposed approach enables the concurrent operation of social learning and reinforcement learning. Specifically, it alternates between a single step of social learning and a single step of MARL, eliminating the need for the time- and computation-intensive two-timescale learning frameworks. Theoretical guarantees are provided to support the effectiveness of the proposed method. Simulation results verify that the performance of the proposed methodology can approach that of reinforcement learning when the true state is known.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy Optimization in Multi-Agent Settings under Partially Observable Environments
Zhaikhan, Ainur
Khammassi, Malek
Sayed, Ali H.
Multiagent Systems
This work leverages adaptive social learning to estimate partially observable global states in multi-agent reinforcement learning (MARL) problems. Unlike existing methods, the proposed approach enables the concurrent operation of social learning and reinforcement learning. Specifically, it alternates between a single step of social learning and a single step of MARL, eliminating the need for the time- and computation-intensive two-timescale learning frameworks. Theoretical guarantees are provided to support the effectiveness of the proposed method. Simulation results verify that the performance of the proposed methodology can approach that of reinforcement learning when the true state is known.
title Policy Optimization in Multi-Agent Settings under Partially Observable Environments
topic Multiagent Systems
url https://arxiv.org/abs/2508.06061