Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Humayoo, Mahammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929599836848128
author Humayoo, Mahammad
author_facet Humayoo, Mahammad
contents Q-Learning is a fundamental off-policy reinforcement learning (RL) algorithm that has the objective of approximating action-value functions in order to learn optimal policies. Nonetheless, it has difficulties in reconciling bias with variance, particularly in the context of long-term rewards. This paper introduces Q($Δ$)-Learning, an extension of TD($Δ$) for the Q-Learning framework. TD($Δ$) facilitates efficient learning over several time scales by breaking the Q($Δ$)-function into distinct discount factors. This approach offers improved learning stability and scalability, especially for long-term tasks where discounting bias may impede convergence. Our methodology guarantees that each element of the Q($Δ$)-function is acquired individually, facilitating expedited convergence on shorter time scales and enhancing the learning of extended time scales. We demonstrate through theoretical analysis and practical evaluations on standard benchmarks like Atari that Q($Δ$)-Learning surpasses conventional Q-Learning and TD learning methods in both tabular and deep RL environments.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14019
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
Humayoo, Mahammad
Machine Learning
F.2.2, I.2.7
Q-Learning is a fundamental off-policy reinforcement learning (RL) algorithm that has the objective of approximating action-value functions in order to learn optimal policies. Nonetheless, it has difficulties in reconciling bias with variance, particularly in the context of long-term rewards. This paper introduces Q($Δ$)-Learning, an extension of TD($Δ$) for the Q-Learning framework. TD($Δ$) facilitates efficient learning over several time scales by breaking the Q($Δ$)-function into distinct discount factors. This approach offers improved learning stability and scalability, especially for long-term tasks where discounting bias may impede convergence. Our methodology guarantees that each element of the Q($Δ$)-function is acquired individually, facilitating expedited convergence on shorter time scales and enhancing the learning of extended time scales. We demonstrate through theoretical analysis and practical evaluations on standard benchmarks like Atari that Q($Δ$)-Learning surpasses conventional Q-Learning and TD learning methods in both tabular and deep RL environments.
title Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
topic Machine Learning
F.2.2, I.2.7
url https://arxiv.org/abs/2411.14019