Decoupled Prioritized Resampling for Offline RL

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yue, Yang, Kang, Bingyi, Ma, Xiao, Yang, Qisen, Huang, Gao, Song, Shiji, Yan, Shuicheng
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910776191614976
author Yue, Yang
Kang, Bingyi
Ma, Xiao
Yang, Qisen
Huang, Gao
Song, Shiji
Yan, Shuicheng
author_facet Yue, Yang
Kang, Bingyi
Ma, Xiao
Yang, Qisen
Huang, Gao
Song, Shiji
Yan, Shuicheng
contents Offline reinforcement learning (RL) is challenged by the distributional shift problem. To address this problem, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. To alleviate this issue, we propose Offline Prioritized Experience Replay (OPER), featuring a class of priority functions designed to prioritize highly-rewarding transitions, making them more frequently visited during training. Through theoretical analysis, we show that this class of priority functions induce an improved behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We develop two practical strategies to obtain priority weights by estimating advantages based on a fitted value network (OPER-A) or utilizing trajectory returns (OPER-R) for quick computation. OPER is a plug-and-play component for offline RL algorithms. As case studies, we evaluate OPER on five different algorithms, including BC, TD3+BC, Onestep RL, CQL, and IQL. Extensive experiments demonstrate that both OPER-A and OPER-R significantly improve the performance for all baseline methods. Codes and priority weights are availiable at https://github.com/sail-sg/OPER.
format Preprint
id arxiv_https___arxiv_org_abs_2306_05412
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Decoupled Prioritized Resampling for Offline RL
Yue, Yang
Kang, Bingyi
Ma, Xiao
Yang, Qisen
Huang, Gao
Song, Shiji
Yan, Shuicheng
Machine Learning
Artificial Intelligence
Offline reinforcement learning (RL) is challenged by the distributional shift problem. To address this problem, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. To alleviate this issue, we propose Offline Prioritized Experience Replay (OPER), featuring a class of priority functions designed to prioritize highly-rewarding transitions, making them more frequently visited during training. Through theoretical analysis, we show that this class of priority functions induce an improved behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We develop two practical strategies to obtain priority weights by estimating advantages based on a fitted value network (OPER-A) or utilizing trajectory returns (OPER-R) for quick computation. OPER is a plug-and-play component for offline RL algorithms. As case studies, we evaluate OPER on five different algorithms, including BC, TD3+BC, Onestep RL, CQL, and IQL. Extensive experiments demonstrate that both OPER-A and OPER-R significantly improve the performance for all baseline methods. Codes and priority weights are availiable at https://github.com/sail-sg/OPER.
title Decoupled Prioritized Resampling for Offline RL
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2306.05412