Compositional Shielding and Reinforcement Learning for Multi-Agent Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Brorholt, Asger Horn, Larsen, Kim Guldstrand, Schilling, Christian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916793621151744
author Brorholt, Asger Horn
Larsen, Kim Guldstrand
Schilling, Christian
author_facet Brorholt, Asger Horn
Larsen, Kim Guldstrand
Schilling, Christian
contents Deep reinforcement learning has emerged as a powerful tool for obtaining high-performance policies. However, the safety of these policies has been a long-standing issue. One promising paradigm to guarantee safety is a shield, which shields a policy from making unsafe actions. However, computing a shield scales exponentially in the number of state variables. This is a particular concern in multi-agent systems with many agents. In this work, we propose a novel approach for multi-agent shielding. We address scalability by computing individual shields for each agent. The challenge is that typical safety specifications are global properties, but the shields of individual agents only ensure local properties. Our key to overcome this challenge is to apply assume-guarantee reasoning. Specifically, we present a sound proof rule that decomposes a (global, complex) safety specification into (local, simple) obligations for the shields of the individual agents. Moreover, we show that applying the shields during reinforcement learning significantly improves the quality of the policies obtained for a given training budget. We demonstrate the effectiveness and scalability of our multi-agent shielding framework in two case studies, reducing the computation time from hours to seconds and achieving fast learning convergence.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10460
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compositional Shielding and Reinforcement Learning for Multi-Agent Systems
Brorholt, Asger Horn
Larsen, Kim Guldstrand
Schilling, Christian
Logic in Computer Science
Artificial Intelligence
Machine Learning
Deep reinforcement learning has emerged as a powerful tool for obtaining high-performance policies. However, the safety of these policies has been a long-standing issue. One promising paradigm to guarantee safety is a shield, which shields a policy from making unsafe actions. However, computing a shield scales exponentially in the number of state variables. This is a particular concern in multi-agent systems with many agents. In this work, we propose a novel approach for multi-agent shielding. We address scalability by computing individual shields for each agent. The challenge is that typical safety specifications are global properties, but the shields of individual agents only ensure local properties. Our key to overcome this challenge is to apply assume-guarantee reasoning. Specifically, we present a sound proof rule that decomposes a (global, complex) safety specification into (local, simple) obligations for the shields of the individual agents. Moreover, we show that applying the shields during reinforcement learning significantly improves the quality of the policies obtained for a given training budget. We demonstrate the effectiveness and scalability of our multi-agent shielding framework in two case studies, reducing the computation time from hours to seconds and achieving fast learning convergence.
title Compositional Shielding and Reinforcement Learning for Multi-Agent Systems
topic Logic in Computer Science
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.10460