Reinforcement Learning with Ensemble Model Predictive Safety Certification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913225090531328 |
|---|---|
| author | Gronauer, Sven Haider, Tom da Roza, Felippe Schmoeller Diepold, Klaus |
| author_facet | Gronauer, Sven Haider, Tom da Roza, Felippe Schmoeller Diepold, Klaus |
| contents | Reinforcement learning algorithms need exploration to learn. However, unsupervised exploration prevents the deployment of such algorithms on safety-critical tasks and limits real-world deployment. In this paper, we propose a new algorithm called Ensemble Model Predictive Safety Certification that combines model-based deep reinforcement learning with tube-based model predictive control to correct the actions taken by a learning agent, keeping safety constraint violations at a minimum through planning. Our approach aims to reduce the amount of prior knowledge about the actual system by requiring only offline data generated by a safe controller. Our results show that we can achieve significantly fewer constraint violations than comparable reinforcement learning methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_04182 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Reinforcement Learning with Ensemble Model Predictive Safety Certification Gronauer, Sven Haider, Tom da Roza, Felippe Schmoeller Diepold, Klaus Machine Learning Robotics Reinforcement learning algorithms need exploration to learn. However, unsupervised exploration prevents the deployment of such algorithms on safety-critical tasks and limits real-world deployment. In this paper, we propose a new algorithm called Ensemble Model Predictive Safety Certification that combines model-based deep reinforcement learning with tube-based model predictive control to correct the actions taken by a learning agent, keeping safety constraint violations at a minimum through planning. Our approach aims to reduce the amount of prior knowledge about the actual system by requiring only offline data generated by a safe controller. Our results show that we can achieve significantly fewer constraint violations than comparable reinforcement learning methods. |
| title | Reinforcement Learning with Ensemble Model Predictive Safety Certification |
| topic | Machine Learning Robotics |
| url | https://arxiv.org/abs/2402.04182 |