Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Koutas, Daniel, Hettegger, Daniel, Papakonstantinou, Kostas G., Straub, Daniel
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916650655154176
author Koutas, Daniel
Hettegger, Daniel
Papakonstantinou, Kostas G.
Straub, Daniel
author_facet Koutas, Daniel
Hettegger, Daniel
Papakonstantinou, Kostas G.
Straub, Daniel
contents We present a novel method for Deep Reinforcement Learning (DRL), incorporating the convex property of the value function over the belief space in Partially Observable Markov Decision Processes (POMDPs). We introduce hard- and soft-enforced convexity as two different approaches, and compare their performance against standard DRL on two well-known POMDP environments, namely the Tiger and FieldVisionRockSample problems. Our findings show that including the convexity feature can substantially increase performance of the agents, as well as increase robustness over the hyperparameter space, especially when testing on out-of-distribution domains. The source code for this work can be found at https://github.com/Dakout/Convex_DRL.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning
Koutas, Daniel
Hettegger, Daniel
Papakonstantinou, Kostas G.
Straub, Daniel
Machine Learning
We present a novel method for Deep Reinforcement Learning (DRL), incorporating the convex property of the value function over the belief space in Partially Observable Markov Decision Processes (POMDPs). We introduce hard- and soft-enforced convexity as two different approaches, and compare their performance against standard DRL on two well-known POMDP environments, namely the Tiger and FieldVisionRockSample problems. Our findings show that including the convexity feature can substantially increase performance of the agents, as well as increase robustness over the hyperparameter space, especially when testing on out-of-distribution domains. The source code for this work can be found at https://github.com/Dakout/Convex_DRL.
title Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2502.09298