Matrix Low-Rank Trust Region Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rozada, Sergio, Marques, Antonio G.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916262760677376
author Rozada, Sergio
Marques, Antonio G.
author_facet Rozada, Sergio
Marques, Antonio G.
contents Most methods in reinforcement learning use a Policy Gradient (PG) approach to learn a parametric stochastic policy that maps states to actions. The standard approach is to implement such a mapping via a neural network (NN) whose parameters are optimized using stochastic gradient descent. However, PG methods are prone to large policy updates that can render learning inefficient. Trust region algorithms, like Trust Region Policy Optimization (TRPO), constrain the policy update step, ensuring monotonic improvements. This paper introduces low-rank matrix-based models as an efficient alternative for estimating the parameters of TRPO algorithms. By gathering the stochastic policy's parameters into a matrix and applying matrix-completion techniques, we promote and enforce low rank. Our numerical studies demonstrate that low-rank matrix-based policy models effectively reduce both computational and sample complexities compared to NN models, while maintaining comparable aggregated rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17625
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Matrix Low-Rank Trust Region Policy Optimization
Rozada, Sergio
Marques, Antonio G.
Machine Learning
Artificial Intelligence
Most methods in reinforcement learning use a Policy Gradient (PG) approach to learn a parametric stochastic policy that maps states to actions. The standard approach is to implement such a mapping via a neural network (NN) whose parameters are optimized using stochastic gradient descent. However, PG methods are prone to large policy updates that can render learning inefficient. Trust region algorithms, like Trust Region Policy Optimization (TRPO), constrain the policy update step, ensuring monotonic improvements. This paper introduces low-rank matrix-based models as an efficient alternative for estimating the parameters of TRPO algorithms. By gathering the stochastic policy's parameters into a matrix and applying matrix-completion techniques, we promote and enforce low rank. Our numerical studies demonstrate that low-rank matrix-based policy models effectively reduce both computational and sample complexities compared to NN models, while maintaining comparable aggregated rewards.
title Matrix Low-Rank Trust Region Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.17625