Difference Rewards Policy Gradients

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Castellini, Jacopo, Devlin, Sam, Oliehoek, Frans A., Savani, Rahul
Format: Preprint
Veröffentlicht: 2020
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910751537496064
author Castellini, Jacopo
Devlin, Sam
Oliehoek, Frans A.
Savani, Rahul
author_facet Castellini, Jacopo
Devlin, Sam
Oliehoek, Frans A.
Savani, Rahul
contents Policy gradient methods have become one of the most popular classes of algorithms for multi-agent reinforcement learning. A key challenge, however, that is not addressed by many of these methods is multi-agent credit assignment: assessing an agent's contribution to the overall performance, which is crucial for learning good policies. We propose a novel algorithm called Dr.Reinforce that explicitly tackles this by combining difference rewards with policy gradients to allow for learning decentralized policies when the reward function is known. By differencing the reward function directly, Dr.Reinforce avoids difficulties associated with learning the Q-function as done by Counterfactual Multiagent Policy Gradients (COMA), a state-of-the-art difference rewards method. For applications where the reward function is unknown, we show the effectiveness of a version of Dr.Reinforce that learns an additional reward network that is used to estimate the difference rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2012_11258
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Difference Rewards Policy Gradients
Castellini, Jacopo
Devlin, Sam
Oliehoek, Frans A.
Savani, Rahul
Multiagent Systems
Machine Learning
I.2.6; I.2.11
Policy gradient methods have become one of the most popular classes of algorithms for multi-agent reinforcement learning. A key challenge, however, that is not addressed by many of these methods is multi-agent credit assignment: assessing an agent's contribution to the overall performance, which is crucial for learning good policies. We propose a novel algorithm called Dr.Reinforce that explicitly tackles this by combining difference rewards with policy gradients to allow for learning decentralized policies when the reward function is known. By differencing the reward function directly, Dr.Reinforce avoids difficulties associated with learning the Q-function as done by Counterfactual Multiagent Policy Gradients (COMA), a state-of-the-art difference rewards method. For applications where the reward function is unknown, we show the effectiveness of a version of Dr.Reinforce that learns an additional reward network that is used to estimate the difference rewards.
title Difference Rewards Policy Gradients
topic Multiagent Systems
Machine Learning
I.2.6; I.2.11
url https://arxiv.org/abs/2012.11258