Heterogeneous Value Decomposition Policy Fusion for Multi-Agent Cooperation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Siying, Zhou, Yang, Zhao, Zhitong, Zhang, Ruoning, Shao, Jinliang, Chen, Wenyu, Cheng, Yuhua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910814807523328
author Wang, Siying
Zhou, Yang
Zhao, Zhitong
Zhang, Ruoning
Shao, Jinliang
Chen, Wenyu
Cheng, Yuhua
author_facet Wang, Siying
Zhou, Yang
Zhao, Zhitong
Zhang, Ruoning
Shao, Jinliang
Chen, Wenyu
Cheng, Yuhua
contents Value decomposition (VD) has become one of the most prominent solutions in cooperative multi-agent reinforcement learning. Most existing methods generally explore how to factorize the joint value and minimize the discrepancies between agent observations and characteristics of environmental states. However, direct decomposition may result in limited representation or difficulty in optimization. Orthogonal to designing a new factorization scheme, in this paper, we propose Heterogeneous Policy Fusion (HPF) to integrate the strengths of various VD methods. We construct a composite policy set to select policies for interaction adaptively. Specifically, this adaptive mechanism allows agents' trajectories to benefit from diverse policy transitions while incorporating the advantages of each factorization method. Additionally, HPF introduces a constraint between these heterogeneous policies to rectify the misleading update caused by the unexpected exploratory or suboptimal non-cooperation. Experimental results on cooperative tasks show HPF's superior performance over multiple baselines, proving its effectiveness and ease of implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Heterogeneous Value Decomposition Policy Fusion for Multi-Agent Cooperation
Wang, Siying
Zhou, Yang
Zhao, Zhitong
Zhang, Ruoning
Shao, Jinliang
Chen, Wenyu
Cheng, Yuhua
Multiagent Systems
Value decomposition (VD) has become one of the most prominent solutions in cooperative multi-agent reinforcement learning. Most existing methods generally explore how to factorize the joint value and minimize the discrepancies between agent observations and characteristics of environmental states. However, direct decomposition may result in limited representation or difficulty in optimization. Orthogonal to designing a new factorization scheme, in this paper, we propose Heterogeneous Policy Fusion (HPF) to integrate the strengths of various VD methods. We construct a composite policy set to select policies for interaction adaptively. Specifically, this adaptive mechanism allows agents' trajectories to benefit from diverse policy transitions while incorporating the advantages of each factorization method. Additionally, HPF introduces a constraint between these heterogeneous policies to rectify the misleading update caused by the unexpected exploratory or suboptimal non-cooperation. Experimental results on cooperative tasks show HPF's superior performance over multiple baselines, proving its effectiveness and ease of implementation.
title Heterogeneous Value Decomposition Policy Fusion for Multi-Agent Cooperation
topic Multiagent Systems
url https://arxiv.org/abs/2502.02875