Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kapoor, Aditya, Swamy, Sushant, Tessera, Kale-ab, Baranwal, Mayank, Sun, Mingfei, Khadilkar, Harshad, Albrecht, Stefano V.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917874131533824
author Kapoor, Aditya
Swamy, Sushant
Tessera, Kale-ab
Baranwal, Mayank
Sun, Mingfei
Khadilkar, Harshad
Albrecht, Stefano V.
author_facet Kapoor, Aditya
Swamy, Sushant
Tessera, Kale-ab
Baranwal, Mayank
Sun, Mingfei
Khadilkar, Harshad
Albrecht, Stefano V.
contents In multi-agent environments, agents often struggle to learn optimal policies due to sparse or delayed global rewards, particularly in long-horizon tasks where it is challenging to evaluate actions at intermediate time steps. We introduce Temporal-Agent Reward Redistribution (TAR$^2$), a novel approach designed to address the agent-temporal credit assignment problem by redistributing sparse rewards both temporally and across agents. TAR$^2$ decomposes sparse global rewards into time-step-specific rewards and calculates agent-specific contributions to these rewards. We theoretically prove that TAR$^2$ is equivalent to potential-based reward shaping, ensuring that the optimal policy remains unchanged. Empirical results demonstrate that TAR$^2$ stabilizes and accelerates the learning process. Additionally, we show that when TAR$^2$ is integrated with single-agent reinforcement learning algorithms, it performs as well as or better than traditional multi-agent reinforcement learning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14779
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning
Kapoor, Aditya
Swamy, Sushant
Tessera, Kale-ab
Baranwal, Mayank
Sun, Mingfei
Khadilkar, Harshad
Albrecht, Stefano V.
Multiagent Systems
Artificial Intelligence
Computer Science and Game Theory
Machine Learning
Robotics
In multi-agent environments, agents often struggle to learn optimal policies due to sparse or delayed global rewards, particularly in long-horizon tasks where it is challenging to evaluate actions at intermediate time steps. We introduce Temporal-Agent Reward Redistribution (TAR$^2$), a novel approach designed to address the agent-temporal credit assignment problem by redistributing sparse rewards both temporally and across agents. TAR$^2$ decomposes sparse global rewards into time-step-specific rewards and calculates agent-specific contributions to these rewards. We theoretically prove that TAR$^2$ is equivalent to potential-based reward shaping, ensuring that the optimal policy remains unchanged. Empirical results demonstrate that TAR$^2$ stabilizes and accelerates the learning process. Additionally, we show that when TAR$^2$ is integrated with single-agent reinforcement learning algorithms, it performs as well as or better than traditional multi-agent reinforcement learning methods.
title Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning
topic Multiagent Systems
Artificial Intelligence
Computer Science and Game Theory
Machine Learning
Robotics
url https://arxiv.org/abs/2412.14779