$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Reddi, Aryaman, Tiboni, Gabriele, Peters, Jan, D'Eramo, Carlo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916950143139840
author Reddi, Aryaman
Tiboni, Gabriele
Peters, Jan
D'Eramo, Carlo
author_facet Reddi, Aryaman
Tiboni, Gabriele
Peters, Jan
D'Eramo, Carlo
contents Actor-critic algorithms for deep multi-agent reinforcement learning (MARL) typically employ a policy update that responds to the current strategies of other agents. While being straightforward, this approach does not account for the updates of other agents at the same update step, resulting in miscoordination. In this paper, we introduce the $K$-Level Policy Gradient (KPG), a method that recursively updates each agent against the updated policies of other agents, speeding up the discovery of effective coordinated policies. We theoretically prove that KPG with finite iterates achieves monotonic convergence to a local Nash equilibrium under certain conditions. We provide principled implementations of KPG by applying it to the deep MARL algorithms MAPPO, MADDPG, and FACMAC. Empirically, we demonstrate superior performance over existing deep MARL algorithms in StarCraft II and multi-agent MuJoCo.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle $K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
Reddi, Aryaman
Tiboni, Gabriele
Peters, Jan
D'Eramo, Carlo
Machine Learning
Artificial Intelligence
Actor-critic algorithms for deep multi-agent reinforcement learning (MARL) typically employ a policy update that responds to the current strategies of other agents. While being straightforward, this approach does not account for the updates of other agents at the same update step, resulting in miscoordination. In this paper, we introduce the $K$-Level Policy Gradient (KPG), a method that recursively updates each agent against the updated policies of other agents, speeding up the discovery of effective coordinated policies. We theoretically prove that KPG with finite iterates achieves monotonic convergence to a local Nash equilibrium under certain conditions. We provide principled implementations of KPG by applying it to the deep MARL algorithms MAPPO, MADDPG, and FACMAC. Empirically, we demonstrate superior performance over existing deep MARL algorithms in StarCraft II and multi-agent MuJoCo.
title $K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.12117