Learning Reward Machines in Cooperative Multi-Agent Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916613207359488 |
|---|---|
| author | Ardon, Leo Furelos-Blanco, Daniel Russo, Alessandra |
| author_facet | Ardon, Leo Furelos-Blanco, Daniel Russo, Alessandra |
| contents | This paper presents a novel approach to Multi-Agent Reinforcement Learning (MARL) that combines cooperative task decomposition with the learning of reward machines (RMs) encoding the structure of the sub-tasks. The proposed method helps deal with the non-Markovian nature of the rewards in partially observable environments and improves the interpretability of the learnt policies required to complete the cooperative task. The RMs associated with each sub-task are learnt in a decentralised manner and then used to guide the behaviour of each agent. By doing so, the complexity of a cooperative multi-agent problem is reduced, allowing for more effective learning. The results suggest that our approach is a promising direction for future research in MARL, especially in complex environments with large state spaces and multiple agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2303_14061 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Learning Reward Machines in Cooperative Multi-Agent Tasks Ardon, Leo Furelos-Blanco, Daniel Russo, Alessandra Artificial Intelligence Multiagent Systems Symbolic Computation This paper presents a novel approach to Multi-Agent Reinforcement Learning (MARL) that combines cooperative task decomposition with the learning of reward machines (RMs) encoding the structure of the sub-tasks. The proposed method helps deal with the non-Markovian nature of the rewards in partially observable environments and improves the interpretability of the learnt policies required to complete the cooperative task. The RMs associated with each sub-task are learnt in a decentralised manner and then used to guide the behaviour of each agent. By doing so, the complexity of a cooperative multi-agent problem is reduced, allowing for more effective learning. The results suggest that our approach is a promising direction for future research in MARL, especially in complex environments with large state spaces and multiple agents. |
| title | Learning Reward Machines in Cooperative Multi-Agent Tasks |
| topic | Artificial Intelligence Multiagent Systems Symbolic Computation |
| url | https://arxiv.org/abs/2303.14061 |