VA-learning as a more efficient alternative to Q-learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Yunhao, Munos, Rémi, Rowland, Mark, Valko, Michal
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917764308926464
author Tang, Yunhao
Munos, Rémi
Rowland, Mark
Valko, Michal
author_facet Tang, Yunhao
Munos, Rémi
Rowland, Mark
Valko, Michal
contents In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function. A natural question is: Why not learn the advantage function directly? In this work, we introduce VA-learning, which directly learns advantage function and value function using bootstrapping, without explicit reference to Q-functions. VA-learning learns off-policy and enjoys similar theoretical guarantees as Q-learning. Thanks to the direct learning of advantage function and value function, VA-learning improves the sample efficiency over Q-learning both in tabular implementations and deep RL agents on Atari-57 games. We also identify a close connection between VA-learning and the dueling architecture, which partially explains why a simple architectural change to DQN agents tends to improve performance.
format Preprint
id arxiv_https___arxiv_org_abs_2305_18161
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VA-learning as a more efficient alternative to Q-learning
Tang, Yunhao
Munos, Rémi
Rowland, Mark
Valko, Michal
Machine Learning
In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function. A natural question is: Why not learn the advantage function directly? In this work, we introduce VA-learning, which directly learns advantage function and value function using bootstrapping, without explicit reference to Q-functions. VA-learning learns off-policy and enjoys similar theoretical guarantees as Q-learning. Thanks to the direct learning of advantage function and value function, VA-learning improves the sample efficiency over Q-learning both in tabular implementations and deep RL agents on Atari-57 games. We also identify a close connection between VA-learning and the dueling architecture, which partially explains why a simple architectural change to DQN agents tends to improve performance.
title VA-learning as a more efficient alternative to Q-learning
topic Machine Learning
url https://arxiv.org/abs/2305.18161