MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Haohan, Cong, Jinmiao, Wang, Shengzhi, Wang, Lu, Liu, Chanjuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911669179908096
author Yu, Haohan
Cong, Jinmiao
Wang, Shengzhi
Wang, Lu
Liu, Chanjuan
author_facet Yu, Haohan
Cong, Jinmiao
Wang, Shengzhi
Wang, Lu
Liu, Chanjuan
contents A key challenge in multi-agent reinforcement learning (MARL) lies in designing learning signals that effectively promote coordination among agents. Designing such signals requires estimating how one agent's current action affects its teammates over future interaction steps. To address this, we introduce Multi-step Advantage-Gated Interventional Causal MARL (MAGIC), a framework that estimates multi-step action effects between agents and selectively converts them into intrinsic rewards. MAGIC uses counterfactual action interventions to compare teammate futures under factual and counterfactual branches, and introduces a gate based on advantage to direct exploration toward beneficial behaviors aligned with the task goal. Experiments on Multi-Agent Particle Environments (MPE) and StarCraft micromanagement benchmarks (SMAC and SMACv2) show that MAGIC consistently outperforms leading prior methods, with average relative final performance improvements of 26.9% and 10.1%, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01805
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning
Yu, Haohan
Cong, Jinmiao
Wang, Shengzhi
Wang, Lu
Liu, Chanjuan
Multiagent Systems
Machine Learning
A key challenge in multi-agent reinforcement learning (MARL) lies in designing learning signals that effectively promote coordination among agents. Designing such signals requires estimating how one agent's current action affects its teammates over future interaction steps. To address this, we introduce Multi-step Advantage-Gated Interventional Causal MARL (MAGIC), a framework that estimates multi-step action effects between agents and selectively converts them into intrinsic rewards. MAGIC uses counterfactual action interventions to compare teammate futures under factual and counterfactual branches, and introduces a gate based on advantage to direct exploration toward beneficial behaviors aligned with the task goal. Experiments on Multi-Agent Particle Environments (MPE) and StarCraft micromanagement benchmarks (SMAC and SMACv2) show that MAGIC consistently outperforms leading prior methods, with average relative final performance improvements of 26.9% and 10.1%, respectively.
title MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning
topic Multiagent Systems
Machine Learning
url https://arxiv.org/abs/2605.01805