Approximation to Deep Q-Network by Stochastic Delay Differential Equations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jianya, Mo, Yingjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915268375085056
author Lu, Jianya
Mo, Yingjun
author_facet Lu, Jianya
Mo, Yingjun
contents Despite the significant breakthroughs that the Deep Q-Network (DQN) has brought to reinforcement learning, its theoretical analysis remains limited. In this paper, we construct a stochastic differential delay equation (SDDE) based on the DQN algorithm and estimate the Wasserstein-1 distance between them. We provide an upper bound for the distance and prove that the distance between the two converges to zero as the step size approaches zero. This result allows us to understand DQN's two key techniques, the experience replay and the target network, from the perspective of continuous systems. Specifically, the delay term in the equation, corresponding to the target network, contributes to the stability of the system. Our approach leverages a refined Lindeberg principle and an operator comparison to establish these results.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00382
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Approximation to Deep Q-Network by Stochastic Delay Differential Equations
Lu, Jianya
Mo, Yingjun
Machine Learning
Probability
Despite the significant breakthroughs that the Deep Q-Network (DQN) has brought to reinforcement learning, its theoretical analysis remains limited. In this paper, we construct a stochastic differential delay equation (SDDE) based on the DQN algorithm and estimate the Wasserstein-1 distance between them. We provide an upper bound for the distance and prove that the distance between the two converges to zero as the step size approaches zero. This result allows us to understand DQN's two key techniques, the experience replay and the target network, from the perspective of continuous systems. Specifically, the delay term in the equation, corresponding to the target network, contributes to the stability of the system. Our approach leverages a refined Lindeberg principle and an operator comparison to establish these results.
title Approximation to Deep Q-Network by Stochastic Delay Differential Equations
topic Machine Learning
Probability
url https://arxiv.org/abs/2505.00382