$β$-DQN: Improving Deep Q-Learning By Evolving the Behavior

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Hongming, Bai, Fengshuo, Xiao, Chenjun, Gao, Chao, Xu, Bo, Müller, Martin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908615687798784
author Zhang, Hongming
Bai, Fengshuo
Xiao, Chenjun
Gao, Chao
Xu, Bo
Müller, Martin
author_facet Zhang, Hongming
Bai, Fengshuo
Xiao, Chenjun
Gao, Chao
Xu, Bo
Müller, Martin
contents While many sophisticated exploration methods have been proposed, their lack of generality and high computational cost often lead researchers to favor simpler methods like $ε$-greedy. Motivated by this, we introduce $β$-DQN, a simple and efficient exploration method that augments the standard DQN with a behavior function $β$. This function estimates the probability that each action has been taken at each state. By leveraging $β$, we generate a population of diverse policies that balance exploration between state-action coverage and overestimation bias correction. An adaptive meta-controller is designed to select an effective policy for each episode, enabling flexible and explainable exploration. $β$-DQN is straightforward to implement and adds minimal computational overhead to the standard DQN. Experiments on both simple and challenging exploration domains show that $β$-DQN outperforms existing baseline methods across a wide range of tasks, providing an effective solution for improving exploration in deep reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle $β$-DQN: Improving Deep Q-Learning By Evolving the Behavior
Zhang, Hongming
Bai, Fengshuo
Xiao, Chenjun
Gao, Chao
Xu, Bo
Müller, Martin
Machine Learning
Artificial Intelligence
While many sophisticated exploration methods have been proposed, their lack of generality and high computational cost often lead researchers to favor simpler methods like $ε$-greedy. Motivated by this, we introduce $β$-DQN, a simple and efficient exploration method that augments the standard DQN with a behavior function $β$. This function estimates the probability that each action has been taken at each state. By leveraging $β$, we generate a population of diverse policies that balance exploration between state-action coverage and overestimation bias correction. An adaptive meta-controller is designed to select an effective policy for each episode, enabling flexible and explainable exploration. $β$-DQN is straightforward to implement and adds minimal computational overhead to the standard DQN. Experiments on both simple and challenging exploration domains show that $β$-DQN outperforms existing baseline methods across a wide range of tasks, providing an effective solution for improving exploration in deep reinforcement learning.
title $β$-DQN: Improving Deep Q-Learning By Evolving the Behavior
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.00913