Tensor and Matrix Low-Rank Value-Function Approximation in Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rozada, Sergio, Paternain, Santiago, Marques, Antonio G.
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910460399321088
author Rozada, Sergio
Paternain, Santiago
Marques, Antonio G.
author_facet Rozada, Sergio
Paternain, Santiago
Marques, Antonio G.
contents Value-function (VF) approximation is a central problem in Reinforcement Learning (RL). Classical non-parametric VF estimation suffers from the curse of dimensionality. As a result, parsimonious parametric models have been adopted to approximate VFs in high-dimensional spaces, with most efforts being focused on linear and neural-network-based approaches. Differently, this paper puts forth a a parsimonious non-parametric approach, where we use stochastic low-rank algorithms to estimate the VF matrix in an online and model-free fashion. Furthermore, as VFs tend to be multi-dimensional, we propose replacing the classical VF matrix representation with a tensor (multi-way array) representation and, then, use the PARAFAC decomposition to design an online model-free tensor low-rank algorithm. Different versions of the algorithms are proposed, their complexity is analyzed, and their performance is assessed numerically using standardized RL environments.
format Preprint
id arxiv_https___arxiv_org_abs_2201_09736
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Tensor and Matrix Low-Rank Value-Function Approximation in Reinforcement Learning
Rozada, Sergio
Paternain, Santiago
Marques, Antonio G.
Machine Learning
Artificial Intelligence
Value-function (VF) approximation is a central problem in Reinforcement Learning (RL). Classical non-parametric VF estimation suffers from the curse of dimensionality. As a result, parsimonious parametric models have been adopted to approximate VFs in high-dimensional spaces, with most efforts being focused on linear and neural-network-based approaches. Differently, this paper puts forth a a parsimonious non-parametric approach, where we use stochastic low-rank algorithms to estimate the VF matrix in an online and model-free fashion. Furthermore, as VFs tend to be multi-dimensional, we propose replacing the classical VF matrix representation with a tensor (multi-way array) representation and, then, use the PARAFAC decomposition to design an online model-free tensor low-rank algorithm. Different versions of the algorithms are proposed, their complexity is analyzed, and their performance is assessed numerically using standardized RL environments.
title Tensor and Matrix Low-Rank Value-Function Approximation in Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2201.09736