Sample-Optimal Zero-Violation Safety For Continuous Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910365515776000 |
|---|---|
| author | Ray, Ritabrata Nakahira, Yorie Kar, Soummya |
| author_facet | Ray, Ritabrata Nakahira, Yorie Kar, Soummya |
| contents | In this paper, we study the problem of ensuring safety with a few shots of samples for partially unknown systems. We first characterize a fundamental limit when producing safe actions is not possible due to insufficient information or samples. Then, we develop a technique that can generate provably safe actions and recovery behaviors using a minimum number of samples. In the performance analysis, we also establish Nagumos theorem - like results with relaxed assumptions, which is potentially useful in other contexts. Finally, we discuss how the proposed method can be integrated into a policy gradient algorithm to assure safety and stability with a handful of samples without stabilizing initial policies or generative models to probe safe actions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_06045 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Sample-Optimal Zero-Violation Safety For Continuous Control Ray, Ritabrata Nakahira, Yorie Kar, Soummya Systems and Control In this paper, we study the problem of ensuring safety with a few shots of samples for partially unknown systems. We first characterize a fundamental limit when producing safe actions is not possible due to insufficient information or samples. Then, we develop a technique that can generate provably safe actions and recovery behaviors using a minimum number of samples. In the performance analysis, we also establish Nagumos theorem - like results with relaxed assumptions, which is potentially useful in other contexts. Finally, we discuss how the proposed method can be integrated into a policy gradient algorithm to assure safety and stability with a handful of samples without stabilizing initial policies or generative models to probe safe actions. |
| title | Sample-Optimal Zero-Violation Safety For Continuous Control |
| topic | Systems and Control |
| url | https://arxiv.org/abs/2403.06045 |