More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishfaq, Haque, Tan, Yixin, Yang, Yu, Lan, Qingfeng, Lu, Jianfeng, Mahmood, A. Rupam, Precup, Doina, Xu, Pan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910492043247616
author Ishfaq, Haque
Tan, Yixin
Yang, Yu
Lan, Qingfeng
Lu, Jianfeng
Mahmood, A. Rupam
Precup, Doina
Xu, Pan
author_facet Ishfaq, Haque
Tan, Yixin
Yang, Yu
Lan, Qingfeng
Lu, Jianfeng
Mahmood, A. Rupam
Precup, Doina
Xu, Pan
contents Thompson sampling (TS) is one of the most popular exploration techniques in reinforcement learning (RL). However, most TS algorithms with theoretical guarantees are difficult to implement and not generalizable to Deep RL. While the emerging approximate sampling-based exploration schemes are promising, most existing algorithms are specific to linear Markov Decision Processes (MDP) with suboptimal regret bounds, or only use the most basic samplers such as Langevin Monte Carlo. In this work, we propose an algorithmic framework that incorporates different approximate sampling methods with the recently proposed Feel-Good Thompson Sampling (FGTS) approach (Zhang, 2022; Dann et al., 2021), which was previously known to be computationally intractable in general. When applied to linear MDPs, our regret analysis yields the best known dependency of regret on dimensionality, surpassing existing randomized algorithms. Additionally, we provide explicit sampling complexity for each employed sampler. Empirically, we show that in tasks where deep exploration is necessary, our proposed algorithms that combine FGTS and approximate sampling perform significantly better compared to other strong baselines. On several challenging games from the Atari 57 suite, our algorithms achieve performance that is either better than or on par with other strong baselines from the deep RL literature.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12241
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
Ishfaq, Haque
Tan, Yixin
Yang, Yu
Lan, Qingfeng
Lu, Jianfeng
Mahmood, A. Rupam
Precup, Doina
Xu, Pan
Machine Learning
Artificial Intelligence
Thompson sampling (TS) is one of the most popular exploration techniques in reinforcement learning (RL). However, most TS algorithms with theoretical guarantees are difficult to implement and not generalizable to Deep RL. While the emerging approximate sampling-based exploration schemes are promising, most existing algorithms are specific to linear Markov Decision Processes (MDP) with suboptimal regret bounds, or only use the most basic samplers such as Langevin Monte Carlo. In this work, we propose an algorithmic framework that incorporates different approximate sampling methods with the recently proposed Feel-Good Thompson Sampling (FGTS) approach (Zhang, 2022; Dann et al., 2021), which was previously known to be computationally intractable in general. When applied to linear MDPs, our regret analysis yields the best known dependency of regret on dimensionality, surpassing existing randomized algorithms. Additionally, we provide explicit sampling complexity for each employed sampler. Empirically, we show that in tasks where deep exploration is necessary, our proposed algorithms that combine FGTS and approximate sampling perform significantly better compared to other strong baselines. On several challenging games from the Atari 57 suite, our algorithms achieve performance that is either better than or on par with other strong baselines from the deep RL literature.
title More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.12241