FOSP: Fine-tuning Offline Safe Policy through World Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Chenyang, Xin, Yucheng, Wu, Silang, He, Longxiang, Yan, Zichen, Tan, Junbo, Wang, Xueqian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915177917579264
author Cao, Chenyang
Xin, Yucheng
Wu, Silang
He, Longxiang
Yan, Zichen
Tan, Junbo
Wang, Xueqian
author_facet Cao, Chenyang
Xin, Yucheng
Wu, Silang
He, Longxiang
Yan, Zichen
Tan, Junbo
Wang, Xueqian
contents Offline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the deployment of vision-based robotic tasks through online fine-tuning an offline pretrained policy. To facilitate effective fine-tuning, we introduce model-based RL, which is known for its data efficiency. Specifically, our method employs in-sample optimization to improve offline training efficiency while incorporating reachability guidance to ensure safety. After obtaining an offline safe policy, a safe policy expansion approach is leveraged for online fine-tuning. The performance of our method is validated on simulation benchmarks with five vision-only tasks and through real-world robot deployment using limited data. It demonstrates that our approach significantly improves the generalization of offline policies to unseen safety-constrained scenarios. To the best of our knowledge, this is the first work to explore offline-to-online RL for safe generalization tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04942
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FOSP: Fine-tuning Offline Safe Policy through World Models
Cao, Chenyang
Xin, Yucheng
Wu, Silang
He, Longxiang
Yan, Zichen
Tan, Junbo
Wang, Xueqian
Robotics
Machine Learning
Offline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the deployment of vision-based robotic tasks through online fine-tuning an offline pretrained policy. To facilitate effective fine-tuning, we introduce model-based RL, which is known for its data efficiency. Specifically, our method employs in-sample optimization to improve offline training efficiency while incorporating reachability guidance to ensure safety. After obtaining an offline safe policy, a safe policy expansion approach is leveraged for online fine-tuning. The performance of our method is validated on simulation benchmarks with five vision-only tasks and through real-world robot deployment using limited data. It demonstrates that our approach significantly improves the generalization of offline policies to unseen safety-constrained scenarios. To the best of our knowledge, this is the first work to explore offline-to-online RL for safe generalization tasks.
title FOSP: Fine-tuning Offline Safe Policy through World Models
topic Robotics
Machine Learning
url https://arxiv.org/abs/2407.04942