Finite-Time Analysis of Temporal Difference Learning with Experience Replay

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lim, Han-Dong, Lee, Donghwan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915242086236160
author Lim, Han-Dong
Lee, Donghwan
author_facet Lim, Han-Dong
Lee, Donghwan
contents Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time behavior, including the finite time bound on mean squared error and sample complexity. On the empirical side, experience replay has been a key ingredient in the success of deep RL algorithms, but its theoretical effects on RL have yet to be fully understood. In this paper, we present a simple decomposition of the Markovian noise terms and provide finite-time error bounds for TD-learning with experience replay. Specifically, under the Markovian observation model, we demonstrate that for both the averaged iterate and final iterate cases, the error term induced by a constant step-size can be effectively controlled by the size of the replay buffer and the mini-batch sampled from the experience replay buffer.
format Preprint
id arxiv_https___arxiv_org_abs_2306_09746
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Finite-Time Analysis of Temporal Difference Learning with Experience Replay
Lim, Han-Dong
Lee, Donghwan
Machine Learning
Artificial Intelligence
Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time behavior, including the finite time bound on mean squared error and sample complexity. On the empirical side, experience replay has been a key ingredient in the success of deep RL algorithms, but its theoretical effects on RL have yet to be fully understood. In this paper, we present a simple decomposition of the Markovian noise terms and provide finite-time error bounds for TD-learning with experience replay. Specifically, under the Markovian observation model, we demonstrate that for both the averaged iterate and final iterate cases, the error term induced by a constant step-size can be effectively controlled by the size of the replay buffer and the mini-batch sampled from the experience replay buffer.
title Finite-Time Analysis of Temporal Difference Learning with Experience Replay
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2306.09746