StaQ it! Growing neural networks for Policy Mirror Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shilova, Alena, Davey, Alex, Driss, Brahim, Akrour, Riad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916796213231616
author Shilova, Alena
Davey, Alex
Driss, Brahim
Akrour, Riad
author_facet Shilova, Alena
Davey, Alex
Driss, Brahim
Akrour, Riad
contents In Reinforcement Learning (RL), regularization has emerged as a popular tool both in theory and practice, typically based either on an entropy bonus or a Kullback-Leibler divergence that constrains successive policies. In practice, these approaches have been shown to improve exploration, robustness and stability, giving rise to popular Deep RL algorithms such as SAC and TRPO. Policy Mirror Descent (PMD) is a theoretical framework that solves this general regularized policy optimization problem, however the closed-form solution involves the sum of all past Q-functions, which is intractable in practice. We propose and analyze PMD-like algorithms that only keep the last $M$ Q-functions in memory, and show that for finite and large enough $M$, a convergent algorithm can be derived, introducing no error in the policy update, unlike prior deep RL PMD implementations. StaQ, the resulting algorithm, enjoys strong theoretical guarantees and is competitive with deep RL baselines, while exhibiting less performance oscillation, paving the way for fully stable deep RL algorithms and providing a testbed for experimentation with Policy Mirror Descent.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StaQ it! Growing neural networks for Policy Mirror Descent
Shilova, Alena
Davey, Alex
Driss, Brahim
Akrour, Riad
Machine Learning
Artificial Intelligence
In Reinforcement Learning (RL), regularization has emerged as a popular tool both in theory and practice, typically based either on an entropy bonus or a Kullback-Leibler divergence that constrains successive policies. In practice, these approaches have been shown to improve exploration, robustness and stability, giving rise to popular Deep RL algorithms such as SAC and TRPO. Policy Mirror Descent (PMD) is a theoretical framework that solves this general regularized policy optimization problem, however the closed-form solution involves the sum of all past Q-functions, which is intractable in practice. We propose and analyze PMD-like algorithms that only keep the last $M$ Q-functions in memory, and show that for finite and large enough $M$, a convergent algorithm can be derived, introducing no error in the policy update, unlike prior deep RL PMD implementations. StaQ, the resulting algorithm, enjoys strong theoretical guarantees and is competitive with deep RL baselines, while exhibiting less performance oscillation, paving the way for fully stable deep RL algorithms and providing a testbed for experimentation with Policy Mirror Descent.
title StaQ it! Growing neural networks for Policy Mirror Descent
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.13862