Bayesian Design Principles for Offline-to-Online Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Hao, Yang, Yiqin, Ye, Jianing, Wu, Chengjie, Mai, Ziqing, Hu, Yujing, Lv, Tangjie, Fan, Changjie, Zhao, Qianchuan, Zhang, Chongjie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913371787362304
author Hu, Hao
Yang, Yiqin
Ye, Jianing
Wu, Chengjie
Mai, Ziqing
Hu, Yujing
Lv, Tangjie
Fan, Changjie
Zhao, Qianchuan
Zhang, Chongjie
author_facet Hu, Hao
Yang, Yiqin
Ye, Jianing
Wu, Chengjie
Mai, Ziqing
Hu, Yujing
Lv, Tangjie
Fan, Changjie
Zhao, Qianchuan
Zhang, Chongjie
contents Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and further online fine-tuning is required. In this paper, we tackle the fundamental dilemma of offline-to-online fine-tuning: if the agent remains pessimistic, it may fail to learn a better policy, while if it becomes optimistic directly, performance may suffer from a sudden drop. We show that Bayesian design principles are crucial in solving such a dilemma. Instead of adopting optimistic or pessimistic policies, the agent should act in a way that matches its belief in optimal policies. Such a probability-matching agent can avoid a sudden performance drop while still being guaranteed to find the optimal policy. Based on our theoretical findings, we introduce a novel algorithm that outperforms existing methods on various benchmarks, demonstrating the efficacy of our approach. Overall, the proposed approach provides a new perspective on offline-to-online RL that has the potential to enable more effective learning from offline data.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20984
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bayesian Design Principles for Offline-to-Online Reinforcement Learning
Hu, Hao
Yang, Yiqin
Ye, Jianing
Wu, Chengjie
Mai, Ziqing
Hu, Yujing
Lv, Tangjie
Fan, Changjie
Zhao, Qianchuan
Zhang, Chongjie
Machine Learning
Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and further online fine-tuning is required. In this paper, we tackle the fundamental dilemma of offline-to-online fine-tuning: if the agent remains pessimistic, it may fail to learn a better policy, while if it becomes optimistic directly, performance may suffer from a sudden drop. We show that Bayesian design principles are crucial in solving such a dilemma. Instead of adopting optimistic or pessimistic policies, the agent should act in a way that matches its belief in optimal policies. Such a probability-matching agent can avoid a sudden performance drop while still being guaranteed to find the optimal policy. Based on our theoretical findings, we introduce a novel algorithm that outperforms existing methods on various benchmarks, demonstrating the efficacy of our approach. Overall, the proposed approach provides a new perspective on offline-to-online RL that has the potential to enable more effective learning from offline data.
title Bayesian Design Principles for Offline-to-Online Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.20984