Gaussian-Mixture-Model Q-Functions for Reinforcement Learning by Riemannian Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vu, Minh, Slavakis, Konstantinos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913495309615104
author Vu, Minh
Slavakis, Konstantinos
author_facet Vu, Minh
Slavakis, Konstantinos
contents This paper establishes a novel role for Gaussian-mixture models (GMMs) as functional approximators of Q-function losses in reinforcement learning (RL). Unlike the existing RL literature, where GMMs play their typical role as estimates of probability density functions, GMMs approximate here Q-function losses. The new Q-function approximators, coined GMM-QFs, are incorporated in Bellman residuals to promote a Riemannian-optimization task as a novel policy-evaluation step in standard policy-iteration schemes. The paper demonstrates how the hyperparameters (means and covariance matrices) of the Gaussian kernels are learned from the data, opening thus the door of RL to the powerful toolbox of Riemannian optimization. Numerical tests show that with no use of experienced data, the proposed design outperforms state-of-the-art methods, even deep Q-networks which use experienced data, on benchmark RL tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04374
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gaussian-Mixture-Model Q-Functions for Reinforcement Learning by Riemannian Optimization
Vu, Minh
Slavakis, Konstantinos
Machine Learning
This paper establishes a novel role for Gaussian-mixture models (GMMs) as functional approximators of Q-function losses in reinforcement learning (RL). Unlike the existing RL literature, where GMMs play their typical role as estimates of probability density functions, GMMs approximate here Q-function losses. The new Q-function approximators, coined GMM-QFs, are incorporated in Bellman residuals to promote a Riemannian-optimization task as a novel policy-evaluation step in standard policy-iteration schemes. The paper demonstrates how the hyperparameters (means and covariance matrices) of the Gaussian kernels are learned from the data, opening thus the door of RL to the powerful toolbox of Riemannian optimization. Numerical tests show that with no use of experienced data, the proposed design outperforms state-of-the-art methods, even deep Q-networks which use experienced data, on benchmark RL tasks.
title Gaussian-Mixture-Model Q-Functions for Reinforcement Learning by Riemannian Optimization
topic Machine Learning
url https://arxiv.org/abs/2409.04374