Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yaodong, Chen, Guangyong, Tang, Hongyao, Liu, Furui, Deng, Danruo, Heng, Pheng Ann
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917911232249856
author Yang, Yaodong
Chen, Guangyong
Tang, Hongyao
Liu, Furui
Deng, Danruo
Heng, Pheng Ann
author_facet Yang, Yaodong
Chen, Guangyong
Tang, Hongyao
Liu, Furui
Deng, Danruo
Heng, Pheng Ann
contents Overestimation in single-agent reinforcement learning has been extensively studied. In contrast, overestimation in the multiagent setting has received comparatively little attention although it increases with the number of agents and leads to severe learning instability. Previous works concentrate on reducing overestimation in the estimation process of target Q-value. They ignore the follow-up optimization process of online Q-network, thus making it hard to fully address the complex multiagent overestimation problem. To solve this challenge, in this study, we first establish an iterative estimation-optimization analysis framework for multiagent value-mixing Q-learning. Our analysis reveals that multiagent overestimation not only comes from the computation of target Q-value but also accumulates in the online Q-network's optimization. Motivated by it, we propose the Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer algorithm to tackle multiagent overestimation from two aspects. First, we extend the random ensemble technique into the estimation of target individual and global Q-values to derive a lower update target. Second, we propose a novel hypernet regularizer on hypernetwork weights and biases to constrain the optimization of online global Q-network to prevent overestimation accumulation. Extensive experiments in MPE and SMAC show that the proposed method successfully addresses overestimation across various tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02018
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
Yang, Yaodong
Chen, Guangyong
Tang, Hongyao
Liu, Furui
Deng, Danruo
Heng, Pheng Ann
Multiagent Systems
Machine Learning
Overestimation in single-agent reinforcement learning has been extensively studied. In contrast, overestimation in the multiagent setting has received comparatively little attention although it increases with the number of agents and leads to severe learning instability. Previous works concentrate on reducing overestimation in the estimation process of target Q-value. They ignore the follow-up optimization process of online Q-network, thus making it hard to fully address the complex multiagent overestimation problem. To solve this challenge, in this study, we first establish an iterative estimation-optimization analysis framework for multiagent value-mixing Q-learning. Our analysis reveals that multiagent overestimation not only comes from the computation of target Q-value but also accumulates in the online Q-network's optimization. Motivated by it, we propose the Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer algorithm to tackle multiagent overestimation from two aspects. First, we extend the random ensemble technique into the estimation of target individual and global Q-values to derive a lower update target. Second, we propose a novel hypernet regularizer on hypernetwork weights and biases to constrain the optimization of online global Q-network to prevent overestimation accumulation. Extensive experiments in MPE and SMAC show that the proposed method successfully addresses overestimation across various tasks.
title Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
topic Multiagent Systems
Machine Learning
url https://arxiv.org/abs/2502.02018