Using Offline Data to Speed Up Reinforcement Learning in Procedurally Generated Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Andres, Alain, Schäfer, Lukas, Albrecht, Stefano V., Del Ser, Javier
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909418592927744
author Andres, Alain
Schäfer, Lukas
Albrecht, Stefano V.
Del Ser, Javier
author_facet Andres, Alain
Schäfer, Lukas
Albrecht, Stefano V.
Del Ser, Javier
contents One of the key challenges of Reinforcement Learning (RL) is the ability of agents to generalise their learned policy to unseen settings. Moreover, training RL agents requires large numbers of interactions with the environment. Motivated by the recent success of Offline RL and Imitation Learning (IL), we conduct a study to investigate whether agents can leverage offline data in the form of trajectories to improve the sample-efficiency in procedurally generated environments. We consider two settings of using IL from offline data for RL: (1) pre-training a policy before online RL training and (2) concurrently training a policy with online RL and IL from offline data. We analyse the impact of the quality (optimality of trajectories) and diversity (number of trajectories and covered level) of available offline trajectories on the effectiveness of both approaches. Across four well-known sparse reward tasks in the MiniGrid environment, we find that using IL for pre-training and concurrently during online RL training both consistently improve the sample-efficiency while converging to optimal policies. Furthermore, we show that pre-training a policy from as few as two trajectories can make the difference between learning an optimal policy at the end of online training and not learning at all. Our findings motivate the widespread adoption of IL for pre-training and concurrent IL in procedurally generated environments whenever offline trajectories are available or can be generated.
format Preprint
id arxiv_https___arxiv_org_abs_2304_09825
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Using Offline Data to Speed Up Reinforcement Learning in Procedurally Generated Environments
Andres, Alain
Schäfer, Lukas
Albrecht, Stefano V.
Del Ser, Javier
Machine Learning
Artificial Intelligence
One of the key challenges of Reinforcement Learning (RL) is the ability of agents to generalise their learned policy to unseen settings. Moreover, training RL agents requires large numbers of interactions with the environment. Motivated by the recent success of Offline RL and Imitation Learning (IL), we conduct a study to investigate whether agents can leverage offline data in the form of trajectories to improve the sample-efficiency in procedurally generated environments. We consider two settings of using IL from offline data for RL: (1) pre-training a policy before online RL training and (2) concurrently training a policy with online RL and IL from offline data. We analyse the impact of the quality (optimality of trajectories) and diversity (number of trajectories and covered level) of available offline trajectories on the effectiveness of both approaches. Across four well-known sparse reward tasks in the MiniGrid environment, we find that using IL for pre-training and concurrently during online RL training both consistently improve the sample-efficiency while converging to optimal policies. Furthermore, we show that pre-training a policy from as few as two trajectories can make the difference between learning an optimal policy at the end of online training and not learning at all. Our findings motivate the widespread adoption of IL for pre-training and concurrent IL in procedurally generated environments whenever offline trajectories are available or can be generated.
title Using Offline Data to Speed Up Reinforcement Learning in Procedurally Generated Environments
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2304.09825