CPIG: Leveraging Consistency Policy with Intention Guidance for Multi-agent Exploration

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fu, Yuqian, Zhu, Yuanheng, Li, Haoran, Zhao, Zijie, Chai, Jiajun, Zhao, Dongbin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929617883889664
author Fu, Yuqian
Zhu, Yuanheng
Li, Haoran
Zhao, Zijie
Chai, Jiajun
Zhao, Dongbin
author_facet Fu, Yuqian
Zhu, Yuanheng
Li, Haoran
Zhao, Zijie
Chai, Jiajun
Zhao, Dongbin
contents Efficient exploration is crucial in cooperative multi-agent reinforcement learning (MARL), especially in sparse-reward settings. However, due to the reliance on the unimodal policy, existing methods are prone to falling into the local optima, hindering the effective exploration of better policies. Furthermore, in sparse-reward settings, each agent tends to receive a scarce reward, which poses significant challenges to inter-agent cooperation. This not only increases the difficulty of policy learning but also degrades the overall performance of multi-agent tasks. To address these issues, we propose a Consistency Policy with Intention Guidance (CPIG), with two primary components: (a) introducing a multimodal policy to enhance the agent's exploration capability, and (b) sharing the intention among agents to foster agent cooperation. For component (a), CPIG incorporates a Consistency model as the policy, leveraging its multimodal nature and stochastic characteristics to facilitate exploration. Regarding component (b), we introduce an Intention Learner to deduce the intention on the global state from each agent's local observation. This intention then serves as a guidance for the Consistency Policy, promoting cooperation among agents. The proposed method is evaluated in multi-agent particle environments (MPE) and multi-agent MuJoCo (MAMuJoCo). Empirical results demonstrate that our method not only achieves comparable performance to various baselines in dense-reward environments but also significantly enhances performance in sparse-reward settings, outperforming state-of-the-art (SOTA) algorithms by 20%.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03603
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CPIG: Leveraging Consistency Policy with Intention Guidance for Multi-agent Exploration
Fu, Yuqian
Zhu, Yuanheng
Li, Haoran
Zhao, Zijie
Chai, Jiajun
Zhao, Dongbin
Multiagent Systems
Efficient exploration is crucial in cooperative multi-agent reinforcement learning (MARL), especially in sparse-reward settings. However, due to the reliance on the unimodal policy, existing methods are prone to falling into the local optima, hindering the effective exploration of better policies. Furthermore, in sparse-reward settings, each agent tends to receive a scarce reward, which poses significant challenges to inter-agent cooperation. This not only increases the difficulty of policy learning but also degrades the overall performance of multi-agent tasks. To address these issues, we propose a Consistency Policy with Intention Guidance (CPIG), with two primary components: (a) introducing a multimodal policy to enhance the agent's exploration capability, and (b) sharing the intention among agents to foster agent cooperation. For component (a), CPIG incorporates a Consistency model as the policy, leveraging its multimodal nature and stochastic characteristics to facilitate exploration. Regarding component (b), we introduce an Intention Learner to deduce the intention on the global state from each agent's local observation. This intention then serves as a guidance for the Consistency Policy, promoting cooperation among agents. The proposed method is evaluated in multi-agent particle environments (MPE) and multi-agent MuJoCo (MAMuJoCo). Empirical results demonstrate that our method not only achieves comparable performance to various baselines in dense-reward environments but also significantly enhances performance in sparse-reward settings, outperforming state-of-the-art (SOTA) algorithms by 20%.
title CPIG: Leveraging Consistency Policy with Intention Guidance for Multi-agent Exploration
topic Multiagent Systems
url https://arxiv.org/abs/2411.03603