Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhuofan, Bollig, Benedikt, Függer, Matthias, Nowak, Thomas, Dréau, Vincent Le
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912539598651392
author Xu, Zhuofan
Bollig, Benedikt
Függer, Matthias
Nowak, Thomas
Dréau, Vincent Le
author_facet Xu, Zhuofan
Bollig, Benedikt
Függer, Matthias
Nowak, Thomas
Dréau, Vincent Le
contents The Centralized Training with Decentralized Execution (CTDE) paradigm has gained significant attention in multi-agent reinforcement learning (MARL) and is the foundation of many recent algorithms. However, decentralized policies operate under partial observability and often yield suboptimal performance compared to centralized policies, while fully centralized approaches typically face scalability challenges as the number of agents increases. We propose Centralized Permutation Equivariant (CPE) learning, a centralized training and execution framework that employs a fully centralized policy to overcome these limitations. Our approach leverages a novel permutation equivariant architecture, Global-Local Permutation Equivariant (GLPE) networks, that is lightweight, scalable, and easy to implement. Experiments show that CPE integrates seamlessly with both value decomposition and actor-critic methods, substantially improving the performance of standard CTDE algorithms across cooperative benchmarks including MPE, SMAC, and RWARE, and matching the performance of state-of-the-art RWARE implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning
Xu, Zhuofan
Bollig, Benedikt
Függer, Matthias
Nowak, Thomas
Dréau, Vincent Le
Multiagent Systems
Artificial Intelligence
Machine Learning
The Centralized Training with Decentralized Execution (CTDE) paradigm has gained significant attention in multi-agent reinforcement learning (MARL) and is the foundation of many recent algorithms. However, decentralized policies operate under partial observability and often yield suboptimal performance compared to centralized policies, while fully centralized approaches typically face scalability challenges as the number of agents increases. We propose Centralized Permutation Equivariant (CPE) learning, a centralized training and execution framework that employs a fully centralized policy to overcome these limitations. Our approach leverages a novel permutation equivariant architecture, Global-Local Permutation Equivariant (GLPE) networks, that is lightweight, scalable, and easy to implement. Experiments show that CPE integrates seamlessly with both value decomposition and actor-critic methods, substantially improving the performance of standard CTDE algorithms across cooperative benchmarks including MPE, SMAC, and RWARE, and matching the performance of state-of-the-art RWARE implementations.
title Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning
topic Multiagent Systems
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.11706