Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Yingru, Xu, Jiawei, Han, Lei, Luo, Zhi-Quan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917693842522112
author Li, Yingru
Xu, Jiawei
Han, Lei
Luo, Zhi-Quan
author_facet Li, Yingru
Xu, Jiawei
Han, Lei
Luo, Zhi-Quan
contents We propose HyperAgent, a reinforcement learning (RL) algorithm based on the hypermodel framework for exploration in RL. HyperAgent allows for the efficient incremental approximation of posteriors associated with an optimal action-value function ($Q^\star$) without the need for conjugacy and follows the greedy policies w.r.t. these approximate posterior samples. We demonstrate that HyperAgent offers robust performance in large-scale deep RL benchmarks. It can solve Deep Sea hard exploration problems with episodes that optimally scale with problem size and exhibits significant efficiency gains in the Atari suite. Implementing HyperAgent requires minimal code addition to well-established deep RL frameworks like DQN. We theoretically prove that, under tabular assumptions, HyperAgent achieves logarithmic per-step computational complexity while attaining sublinear regret, matching the best known randomized tabular RL algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10228
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
Li, Yingru
Xu, Jiawei
Han, Lei
Luo, Zhi-Quan
Machine Learning
Artificial Intelligence
We propose HyperAgent, a reinforcement learning (RL) algorithm based on the hypermodel framework for exploration in RL. HyperAgent allows for the efficient incremental approximation of posteriors associated with an optimal action-value function ($Q^\star$) without the need for conjugacy and follows the greedy policies w.r.t. these approximate posterior samples. We demonstrate that HyperAgent offers robust performance in large-scale deep RL benchmarks. It can solve Deep Sea hard exploration problems with episodes that optimally scale with problem size and exhibits significant efficiency gains in the Atari suite. Implementing HyperAgent requires minimal code addition to well-established deep RL frameworks like DQN. We theoretically prove that, under tabular assumptions, HyperAgent achieves logarithmic per-step computational complexity while attaining sublinear regret, matching the best known randomized tabular RL algorithm.
title Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.10228