On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Enami, Shoju, Kashima, Kenji
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915877050384384
author Enami, Shoju
Kashima, Kenji
author_facet Enami, Shoju
Kashima, Kenji
contents In recent years, mutual information optimal control has been proposed as an extension of maximum entropy optimal control. Both approaches introduce regularization terms to render the policy stochastic, and it is important to theoretically clarify the relationship between the temperature parameter (i.e., the coefficient of the regularization term) and the stochasticity of the policy. Unlike in maximum entropy optimal control, this relationship remains unexplored in mutual information optimal control. In this paper, we investigate this relationship for a mutual information optimal control problem (MIOCP) of discrete-time linear systems. After extending the result of a previous study of the MIOCP, we establish the existence of an optimal policy of the MIOCP, and then derive the respective conditions on the temperature parameter under which the optimal policy becomes stochastic and deterministic. Furthermore, we also derive the respective conditions on the temperature parameter under which the policy obtained by an alternating optimization algorithm becomes stochastic and deterministic. The validity of the theoretical results is demonstrated through numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
Enami, Shoju
Kashima, Kenji
Optimization and Control
Machine Learning
Systems and Control
In recent years, mutual information optimal control has been proposed as an extension of maximum entropy optimal control. Both approaches introduce regularization terms to render the policy stochastic, and it is important to theoretically clarify the relationship between the temperature parameter (i.e., the coefficient of the regularization term) and the stochasticity of the policy. Unlike in maximum entropy optimal control, this relationship remains unexplored in mutual information optimal control. In this paper, we investigate this relationship for a mutual information optimal control problem (MIOCP) of discrete-time linear systems. After extending the result of a previous study of the MIOCP, we establish the existence of an optimal policy of the MIOCP, and then derive the respective conditions on the temperature parameter under which the optimal policy becomes stochastic and deterministic. Furthermore, we also derive the respective conditions on the temperature parameter under which the policy obtained by an alternating optimization algorithm becomes stochastic and deterministic. The validity of the theoretical results is demonstrated through numerical experiments.
title On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
topic Optimization and Control
Machine Learning
Systems and Control
url https://arxiv.org/abs/2507.21543