Mirror descent for constrained stochastic control problems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sethi, Deven, Šiška, David
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913872163635200
author Sethi, Deven
Šiška, David
author_facet Sethi, Deven
Šiška, David
contents Mirror descent is a well established tool for solving convex optimization problems with convex constraints. This article introduces continuous-time mirror descent dynamics for approximating optimal Markov controls for stochastic control problems with the action space being bounded and convex. We show that if the Hamiltonian is uniformly convex in its action variable then mirror descent converges linearly while if it is uniformly strongly convex relative to an appropriate Bregman divergence, then the mirror flow converges exponentially. The two fundamental difficulties that must be overcome to prove such results are: first, the inherent lack of convexity of the map from Markov controls to the corresponding value function. Second, maintaining sufficient regularity of the value function and the Markov controls along the mirror descent updates. The first issue is handled using the performance difference lemma, while the second using careful Sobolev space estimates for the solutions of the associated linear PDEs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mirror descent for constrained stochastic control problems
Sethi, Deven
Šiška, David
Optimization and Control
Mirror descent is a well established tool for solving convex optimization problems with convex constraints. This article introduces continuous-time mirror descent dynamics for approximating optimal Markov controls for stochastic control problems with the action space being bounded and convex. We show that if the Hamiltonian is uniformly convex in its action variable then mirror descent converges linearly while if it is uniformly strongly convex relative to an appropriate Bregman divergence, then the mirror flow converges exponentially. The two fundamental difficulties that must be overcome to prove such results are: first, the inherent lack of convexity of the map from Markov controls to the corresponding value function. Second, maintaining sufficient regularity of the value function and the Markov controls along the mirror descent updates. The first issue is handled using the performance difference lemma, while the second using careful Sobolev space estimates for the solutions of the associated linear PDEs.
title Mirror descent for constrained stochastic control problems
topic Optimization and Control
url https://arxiv.org/abs/2506.02564