Counterfactual Multi-Agent Policy Gradients
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2017
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913606325501952 |
|---|---|
| author | Foerster, Jakob Farquhar, Gregory Afouras, Triantafyllos Nardelli, Nantas Whiteson, Shimon |
| author_facet | Foerster, Jakob Farquhar, Gregory Afouras, Triantafyllos Nardelli, Nantas Whiteson, Shimon |
| contents | Cooperative multi-agent systems can be naturally used to model many real world problems, such as network packet routing and the coordination of autonomous vehicles. There is a great need for new reinforcement learning methods that can efficiently learn decentralised policies for such systems. To this end, we propose a new multi-agent actor-critic method called counterfactual multi-agent (COMA) policy gradients. COMA uses a centralised critic to estimate the Q-function and decentralised actors to optimise the agents' policies. In addition, to address the challenges of multi-agent credit assignment, it uses a counterfactual baseline that marginalises out a single agent's action, while keeping the other agents' actions fixed. COMA also uses a critic representation that allows the counterfactual baseline to be computed efficiently in a single forward pass. We evaluate COMA in the testbed of StarCraft unit micromanagement, using a decentralised variant with significant partial observability. COMA significantly improves average performance over other multi-agent actor-critic methods in this setting, and the best performing agents are competitive with state-of-the-art centralised controllers that get access to the full state. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_1705_08926 |
| institution | arXiv |
| publishDate | 2017 |
| record_format | arxiv |
| spellingShingle | Counterfactual Multi-Agent Policy Gradients Foerster, Jakob Farquhar, Gregory Afouras, Triantafyllos Nardelli, Nantas Whiteson, Shimon Artificial Intelligence Multiagent Systems Cooperative multi-agent systems can be naturally used to model many real world problems, such as network packet routing and the coordination of autonomous vehicles. There is a great need for new reinforcement learning methods that can efficiently learn decentralised policies for such systems. To this end, we propose a new multi-agent actor-critic method called counterfactual multi-agent (COMA) policy gradients. COMA uses a centralised critic to estimate the Q-function and decentralised actors to optimise the agents' policies. In addition, to address the challenges of multi-agent credit assignment, it uses a counterfactual baseline that marginalises out a single agent's action, while keeping the other agents' actions fixed. COMA also uses a critic representation that allows the counterfactual baseline to be computed efficiently in a single forward pass. We evaluate COMA in the testbed of StarCraft unit micromanagement, using a decentralised variant with significant partial observability. COMA significantly improves average performance over other multi-agent actor-critic methods in this setting, and the best performing agents are competitive with state-of-the-art centralised controllers that get access to the full state. |
| title | Counterfactual Multi-Agent Policy Gradients |
| topic | Artificial Intelligence Multiagent Systems |
| url | https://arxiv.org/abs/1705.08926 |