Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kapoor, Aditya, Tessera, Kale-ab, Baranwal, Mayank, Khadilkar, Harshad, Peters, Jan, Albrecht, Stefano, Sun, Mingfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914120086847488
author Kapoor, Aditya
Tessera, Kale-ab
Baranwal, Mayank
Khadilkar, Harshad
Peters, Jan
Albrecht, Stefano
Sun, Mingfei
author_facet Kapoor, Aditya
Tessera, Kale-ab
Baranwal, Mayank
Khadilkar, Harshad
Peters, Jan
Albrecht, Stefano
Sun, Mingfei
contents Credit assignmen, disentangling each agent's contribution to a shared reward, is a critical challenge in cooperative multi-agent reinforcement learning (MARL). To be effective, credit assignment methods must preserve the environment's optimal policy. Some recent approaches attempt this by enforcing return equivalence, where the sum of distributed rewards must equal the team reward. However, their guarantees are conditional on a learned model's regression accuracy, making them unreliable in practice. We introduce Temporal-Agent Reward Redistribution (TAR$^2$), an approach that decouples credit modeling from this constraint. A neural network learns unnormalized contribution scores, while a separate, deterministic normalization step enforces return equivalence by construction. We demonstrate that this method is equivalent to a valid Potential-Based Reward Shaping (PBRS), which guarantees the optimal policy is preserved regardless of model accuracy. Empirically, on challenging SMACLite and Google Research Football (GRF) benchmarks, TAR$^2$ accelerates learning and achieves higher final performance than strong baselines. These results establish our method as an effective solution for the agent-temporal credit assignment problem.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
Kapoor, Aditya
Tessera, Kale-ab
Baranwal, Mayank
Khadilkar, Harshad
Peters, Jan
Albrecht, Stefano
Sun, Mingfei
Multiagent Systems
Artificial Intelligence
Machine Learning
Robotics
Credit assignmen, disentangling each agent's contribution to a shared reward, is a critical challenge in cooperative multi-agent reinforcement learning (MARL). To be effective, credit assignment methods must preserve the environment's optimal policy. Some recent approaches attempt this by enforcing return equivalence, where the sum of distributed rewards must equal the team reward. However, their guarantees are conditional on a learned model's regression accuracy, making them unreliable in practice. We introduce Temporal-Agent Reward Redistribution (TAR$^2$), an approach that decouples credit modeling from this constraint. A neural network learns unnormalized contribution scores, while a separate, deterministic normalization step enforces return equivalence by construction. We demonstrate that this method is equivalent to a valid Potential-Based Reward Shaping (PBRS), which guarantees the optimal policy is preserved regardless of model accuracy. Empirically, on challenging SMACLite and Google Research Football (GRF) benchmarks, TAR$^2$ accelerates learning and achieves higher final performance than strong baselines. These results establish our method as an effective solution for the agent-temporal credit assignment problem.
title Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
topic Multiagent Systems
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2502.04864