A learning-based approach to stochastic optimal control under reach-avoid constraint

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ni, Tingting, Kamgarpour, Maryam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908562378194944
author Ni, Tingting
Kamgarpour, Maryam
author_facet Ni, Tingting
Kamgarpour, Maryam
contents We develop a model-free approach to optimally control stochastic, Markovian systems subject to a reach-avoid constraint. Specifically, the state trajectory must remain within a safe set while reaching a target set within a finite time horizon. Due to the time-dependent nature of these constraints, we show that, in general, the optimal policy for this constrained stochastic control problem is non-Markovian, which increases the computational complexity. To address this challenge, we apply the state-augmentation technique from arXiv:2402.19360, reformulating the problem as a constrained Markov decision process (CMDP) on an extended state space. This transformation allows us to search for a Markovian policy, avoiding the complexity of non-Markovian policies. To learn the optimal policy without a system model, and using only trajectory data, we develop a log-barrier policy gradient approach. We prove that under suitable assumptions, the policy parameters converge to the optimal parameters, while ensuring that the system trajectories satisfy the stochastic reach-avoid constraint with high probability.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16561
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A learning-based approach to stochastic optimal control under reach-avoid constraint
Ni, Tingting
Kamgarpour, Maryam
Optimization and Control
Machine Learning
We develop a model-free approach to optimally control stochastic, Markovian systems subject to a reach-avoid constraint. Specifically, the state trajectory must remain within a safe set while reaching a target set within a finite time horizon. Due to the time-dependent nature of these constraints, we show that, in general, the optimal policy for this constrained stochastic control problem is non-Markovian, which increases the computational complexity. To address this challenge, we apply the state-augmentation technique from arXiv:2402.19360, reformulating the problem as a constrained Markov decision process (CMDP) on an extended state space. This transformation allows us to search for a Markovian policy, avoiding the complexity of non-Markovian policies. To learn the optimal policy without a system model, and using only trajectory data, we develop a log-barrier policy gradient approach. We prove that under suitable assumptions, the policy parameters converge to the optimal parameters, while ensuring that the system trajectories satisfy the stochastic reach-avoid constraint with high probability.
title A learning-based approach to stochastic optimal control under reach-avoid constraint
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2412.16561