H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, Haoyi, Ji, Tianying, Liu, Bingqi, Zhao, Haocheng, Zhu, Xiangyu, Zheng, Jianying, Huang, Pengfei, Zhou, Guyue, Hu, Jianming, Zhan, Xianyuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909580628328448
author Niu, Haoyi
Ji, Tianying
Liu, Bingqi
Zhao, Haocheng
Zhu, Xiangyu
Zheng, Jianying
Huang, Pengfei
Zhou, Guyue
Hu, Jianming
Zhan, Xianyuan
author_facet Niu, Haoyi
Ji, Tianying
Liu, Bingqi
Zhao, Haocheng
Zhu, Xiangyu
Zheng, Jianying
Huang, Pengfei
Zhou, Guyue
Hu, Jianming
Zhan, Xianyuan
contents Solving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called H2O+, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environment. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility over advanced cross-domain online and offline RL algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12716
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
Niu, Haoyi
Ji, Tianying
Liu, Bingqi
Zhao, Haocheng
Zhu, Xiangyu
Zheng, Jianying
Huang, Pengfei
Zhou, Guyue
Hu, Jianming
Zhan, Xianyuan
Machine Learning
Artificial Intelligence
Robotics
Solving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called H2O+, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environment. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility over advanced cross-domain online and offline RL algorithms.
title H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2309.12716