ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Wen-Tse, Li, Yuxuan, Huang, Shiyu, Chen, Jiayu, Schneider, Jeff
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912872324399104
author Chen, Wen-Tse
Li, Yuxuan
Huang, Shiyu
Chen, Jiayu
Schneider, Jeff
author_facet Chen, Wen-Tse
Li, Yuxuan
Huang, Shiyu
Chen, Jiayu
Schneider, Jeff
contents Multi-agent credit assignment is a fundamental challenge for cooperative multi-agent reinforcement learning (MARL), where a team of agents learn from shared reward signals. The Individual-Global-Max (IGM) condition is a widely used principle for multi-agent credit assignment, requiring that the joint action determined by individual Q-functions maximizes the global Q-value. Meanwhile, the principle of maximum entropy has been leveraged to enhance exploration in MARL. However, we identify a critical limitation in existing maximum entropy MARL methods: a misalignment arises between local policies and the joint policy that maximizes the global Q-value, leading to violations of the IGM condition. To address this misalignment, we propose an order-preserving transformation. Building on it, we introduce ME-IGM, a novel maximum entropy MARL algorithm compatible with any credit assignment mechanism that satisfies the IGM condition while enjoying the benefits of maximum entropy exploration. We empirically evaluate two variants of ME-IGM: ME-QMIX and ME-QPLEX, in non-monotonic matrix games, and demonstrate their state-of-the-art performance across 17 scenarios in SMAC-v2 and Overcooked.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
Chen, Wen-Tse
Li, Yuxuan
Huang, Shiyu
Chen, Jiayu
Schneider, Jeff
Machine Learning
I.2.11; I.2.6
Multi-agent credit assignment is a fundamental challenge for cooperative multi-agent reinforcement learning (MARL), where a team of agents learn from shared reward signals. The Individual-Global-Max (IGM) condition is a widely used principle for multi-agent credit assignment, requiring that the joint action determined by individual Q-functions maximizes the global Q-value. Meanwhile, the principle of maximum entropy has been leveraged to enhance exploration in MARL. However, we identify a critical limitation in existing maximum entropy MARL methods: a misalignment arises between local policies and the joint policy that maximizes the global Q-value, leading to violations of the IGM condition. To address this misalignment, we propose an order-preserving transformation. Building on it, we introduce ME-IGM, a novel maximum entropy MARL algorithm compatible with any credit assignment mechanism that satisfies the IGM condition while enjoying the benefits of maximum entropy exploration. We empirically evaluate two variants of ME-IGM: ME-QMIX and ME-QPLEX, in non-monotonic matrix games, and demonstrate their state-of-the-art performance across 17 scenarios in SMAC-v2 and Overcooked.
title ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
topic Machine Learning
I.2.11; I.2.6
url https://arxiv.org/abs/2406.13930