Saved in:
Bibliographic Details
Main Authors: Hudák, David, Galesloot, Maris F. L., Tappler, Martin, Kurečka, Martin, Jansen, Nils, Češka, Milan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.08734
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915904982351872
author Hudák, David
Galesloot, Maris F. L.
Tappler, Martin
Kurečka, Martin
Jansen, Nils
Češka, Milan
author_facet Hudák, David
Galesloot, Maris F. L.
Tappler, Martin
Kurečka, Martin
Jansen, Nils
Češka, Milan
contents Solving partially observable Markov decision processes (POMDPs) requires computing policies under imperfect state information. Despite recent advances, the scalability of existing POMDP solvers remains limited. Moreover, many settings require a policy that is robust across multiple POMDPs, further aggravating the scalability issue. We propose the Lexpop framework for POMDP solving. Lexpop (1) employs deep reinforcement learning to train a neural policy, represented by a recurrent neural network, and (2) constructs a finite-state controller mimicking the neural policy through efficient extraction methods. Crucially, unlike neural policies, such controllers can be formally evaluated, providing performance guarantees. We extend Lexpop to compute robust policies for hidden-model POMDPs (HM-POMDPs), which describe finite sets of POMDPs. We associate every extracted controller with its worst-case POMDP. Using a set of such POMDPs, we iteratively train a robust neural policy and consequently extract a robust controller. Our experiments show that on problems with large state spaces, Lexpop outperforms state-of-the-art solvers for POMDPs as well as HM-POMDPs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08734
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning
Hudák, David
Galesloot, Maris F. L.
Tappler, Martin
Kurečka, Martin
Jansen, Nils
Češka, Milan
Artificial Intelligence
Solving partially observable Markov decision processes (POMDPs) requires computing policies under imperfect state information. Despite recent advances, the scalability of existing POMDP solvers remains limited. Moreover, many settings require a policy that is robust across multiple POMDPs, further aggravating the scalability issue. We propose the Lexpop framework for POMDP solving. Lexpop (1) employs deep reinforcement learning to train a neural policy, represented by a recurrent neural network, and (2) constructs a finite-state controller mimicking the neural policy through efficient extraction methods. Crucially, unlike neural policies, such controllers can be formally evaluated, providing performance guarantees. We extend Lexpop to compute robust policies for hidden-model POMDPs (HM-POMDPs), which describe finite sets of POMDPs. We associate every extracted controller with its worst-case POMDP. Using a set of such POMDPs, we iteratively train a robust neural policy and consequently extract a robust controller. Our experiments show that on problems with large state spaces, Lexpop outperforms state-of-the-art solvers for POMDPs as well as HM-POMDPs.
title Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2602.08734