MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ziyan, Du, Yali, Zhang, Yudi, Fang, Meng, Huang, Biwei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911739885387776
author Wang, Ziyan
Du, Yali
Zhang, Yudi
Fang, Meng
Huang, Biwei
author_facet Wang, Ziyan
Du, Yali
Zhang, Yudi
Fang, Meng
Huang, Biwei
contents Offline Multi-agent Reinforcement Learning (MARL) is valuable in scenarios where online interaction is impractical or risky. While independent learning in MARL offers flexibility and scalability, accurately assigning credit to individual agents in offline settings poses challenges because interactions with an environment are prohibited. In this paper, we propose a new framework, namely Multi-Agent Causal Credit Assignment (MACCA), to address credit assignment in the offline MARL setting. Our approach, MACCA, characterizing the generative process as a Dynamic Bayesian Network, captures relationships between environmental variables, states, actions, and rewards. Estimating this model on offline data, MACCA can learn each agent's contribution by analyzing the causal relationship of their individual rewards, ensuring accurate and interpretable credit assignment. Additionally, the modularity of our approach allows it to integrate with various offline MARL methods seamlessly. Theoretically, we proved that under the setting of the offline dataset, the underlying causal structure and the function for generating the individual rewards of agents are identifiable, which laid the foundation for the correctness of our modeling. In our experiments, we demonstrate that MACCA not only outperforms state-of-the-art methods but also enhances performance when integrated with other backbones.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03644
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment
Wang, Ziyan
Du, Yali
Zhang, Yudi
Fang, Meng
Huang, Biwei
Machine Learning
Multiagent Systems
Offline Multi-agent Reinforcement Learning (MARL) is valuable in scenarios where online interaction is impractical or risky. While independent learning in MARL offers flexibility and scalability, accurately assigning credit to individual agents in offline settings poses challenges because interactions with an environment are prohibited. In this paper, we propose a new framework, namely Multi-Agent Causal Credit Assignment (MACCA), to address credit assignment in the offline MARL setting. Our approach, MACCA, characterizing the generative process as a Dynamic Bayesian Network, captures relationships between environmental variables, states, actions, and rewards. Estimating this model on offline data, MACCA can learn each agent's contribution by analyzing the causal relationship of their individual rewards, ensuring accurate and interpretable credit assignment. Additionally, the modularity of our approach allows it to integrate with various offline MARL methods seamlessly. Theoretically, we proved that under the setting of the offline dataset, the underlying causal structure and the function for generating the individual rewards of agents are identifiable, which laid the foundation for the correctness of our modeling. In our experiments, we demonstrate that MACCA not only outperforms state-of-the-art methods but also enhances performance when integrated with other backbones.
title MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2312.03644