VA-learning as a more efficient alternative to Q-learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917764308926464 |
|---|---|
| author | Tang, Yunhao Munos, Rémi Rowland, Mark Valko, Michal |
| author_facet | Tang, Yunhao Munos, Rémi Rowland, Mark Valko, Michal |
| contents | In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function. A natural question is: Why not learn the advantage function directly? In this work, we introduce VA-learning, which directly learns advantage function and value function using bootstrapping, without explicit reference to Q-functions. VA-learning learns off-policy and enjoys similar theoretical guarantees as Q-learning. Thanks to the direct learning of advantage function and value function, VA-learning improves the sample efficiency over Q-learning both in tabular implementations and deep RL agents on Atari-57 games. We also identify a close connection between VA-learning and the dueling architecture, which partially explains why a simple architectural change to DQN agents tends to improve performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_18161 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | VA-learning as a more efficient alternative to Q-learning Tang, Yunhao Munos, Rémi Rowland, Mark Valko, Michal Machine Learning In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function. A natural question is: Why not learn the advantage function directly? In this work, we introduce VA-learning, which directly learns advantage function and value function using bootstrapping, without explicit reference to Q-functions. VA-learning learns off-policy and enjoys similar theoretical guarantees as Q-learning. Thanks to the direct learning of advantage function and value function, VA-learning improves the sample efficiency over Q-learning both in tabular implementations and deep RL agents on Atari-57 games. We also identify a close connection between VA-learning and the dueling architecture, which partially explains why a simple architectural change to DQN agents tends to improve performance. |
| title | VA-learning as a more efficient alternative to Q-learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2305.18161 |