Deep Gaussian Process Proximal Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: van der Lende, Matthijs, Cardenas-Cartagena, Juan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915686749569024
author van der Lende, Matthijs
Cardenas-Cartagena, Juan
author_facet van der Lende, Matthijs
Cardenas-Cartagena, Juan
contents Uncertainty estimation for Reinforcement Learning (RL) is a critical component in control tasks where agents must balance safe exploration and efficient learning. While deep neural networks have enabled breakthroughs in RL, they often lack calibrated uncertainty estimates. We introduce Deep Gaussian Process Proximal Policy Optimization (GPPO), a scalable, model-free actor-critic algorithm that leverages Deep Gaussian Processes (DGPs) to approximate both the policy and value function. GPPO maintains competitive performance with respect to Proximal Policy Optimization on standard high-dimensional continuous control benchmarks while providing well-calibrated uncertainty estimates that can inform safer and more effective exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18214
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Gaussian Process Proximal Policy Optimization
van der Lende, Matthijs
Cardenas-Cartagena, Juan
Machine Learning
Uncertainty estimation for Reinforcement Learning (RL) is a critical component in control tasks where agents must balance safe exploration and efficient learning. While deep neural networks have enabled breakthroughs in RL, they often lack calibrated uncertainty estimates. We introduce Deep Gaussian Process Proximal Policy Optimization (GPPO), a scalable, model-free actor-critic algorithm that leverages Deep Gaussian Processes (DGPs) to approximate both the policy and value function. GPPO maintains competitive performance with respect to Proximal Policy Optimization on standard high-dimensional continuous control benchmarks while providing well-calibrated uncertainty estimates that can inform safer and more effective exploration.
title Deep Gaussian Process Proximal Policy Optimization
topic Machine Learning
url https://arxiv.org/abs/2511.18214