Predictable Interval MDPs through Entropy Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: van Zutphen, Menno, Delimpaltadakis, Giannis, Heemels, Maurice, Antunes, Duarte
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910024016592896
author van Zutphen, Menno
Delimpaltadakis, Giannis
Heemels, Maurice
Antunes, Duarte
author_facet van Zutphen, Menno
Delimpaltadakis, Giannis
Heemels, Maurice
Antunes, Duarte
contents Regularization of control policies using entropy can be instrumental in adjusting predictability of real-world systems. Applications benefiting from such approaches range from, e.g., cybersecurity, which aims at maximal unpredictability, to human-robot interaction, where predictable behavior is highly desirable. In this paper, we consider entropy regularization for interval Markov decision processes (IMDPs). IMDPs are uncertain MDPs, where transition probabilities are only known to belong to intervals. Lately, IMDPs have gained significant popularity in the context of abstracting stochastic systems for control design. In this work, we address robust minimization of the linear combination of entropy and a standard cumulative cost in IMDPs, thereby establishing a trade-off between optimality and predictability. We show that optimal deterministic policies exist, and devise a value-iteration algorithm to compute them. The algorithm solves a number of convex programs at each step. Finally, through an illustrative example we show the benefits of penalizing entropy in IMDPs.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16711
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Predictable Interval MDPs through Entropy Regularization
van Zutphen, Menno
Delimpaltadakis, Giannis
Heemels, Maurice
Antunes, Duarte
Systems and Control
Regularization of control policies using entropy can be instrumental in adjusting predictability of real-world systems. Applications benefiting from such approaches range from, e.g., cybersecurity, which aims at maximal unpredictability, to human-robot interaction, where predictable behavior is highly desirable. In this paper, we consider entropy regularization for interval Markov decision processes (IMDPs). IMDPs are uncertain MDPs, where transition probabilities are only known to belong to intervals. Lately, IMDPs have gained significant popularity in the context of abstracting stochastic systems for control design. In this work, we address robust minimization of the linear combination of entropy and a standard cumulative cost in IMDPs, thereby establishing a trade-off between optimality and predictability. We show that optimal deterministic policies exist, and devise a value-iteration algorithm to compute them. The algorithm solves a number of convex programs at each step. Finally, through an illustrative example we show the benefits of penalizing entropy in IMDPs.
title Predictable Interval MDPs through Entropy Regularization
topic Systems and Control
url https://arxiv.org/abs/2403.16711