Actively Learning Reinforcement Learning: A Stochastic Optimal Control Approach

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ramadan, Mohammad S., Hayajnh, Mahmoud A., Tolley, Michael T., Vamvoudakis, Kyriakos G.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910593811742720
author Ramadan, Mohammad S.
Hayajnh, Mahmoud A.
Tolley, Michael T.
Vamvoudakis, Kyriakos G.
author_facet Ramadan, Mohammad S.
Hayajnh, Mahmoud A.
Tolley, Michael T.
Vamvoudakis, Kyriakos G.
contents In this paper we propose a framework towards achieving two intertwined objectives: (i) equipping reinforcement learning with active exploration and deliberate information gathering, such that it regulates state and parameter uncertainties resulting from modeling mismatches and noisy sensory; and (ii) overcoming the computational intractability of stochastic optimal control. We approach both objectives by using reinforcement learning to compute the stochastic optimal control law. On one hand, we avoid the curse of dimensionality prohibiting the direct solution of the stochastic dynamic programming equation. On the other hand, the resulting stochastic optimal control reinforcement learning agent admits caution and probing, that is, optimal online exploration and exploitation. Unlike fixed exploration and exploitation balance, caution and probing are employed automatically by the controller in real-time, even after the learning process is terminated. We conclude the paper with a numerical simulation, illustrating how a Linear Quadratic Regulator with the certainty equivalence assumption may lead to poor performance and filter divergence, while our proposed approach is stabilizing, of an acceptable performance, and computationally convenient.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10831
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Actively Learning Reinforcement Learning: A Stochastic Optimal Control Approach
Ramadan, Mohammad S.
Hayajnh, Mahmoud A.
Tolley, Michael T.
Vamvoudakis, Kyriakos G.
Machine Learning
Systems and Control
In this paper we propose a framework towards achieving two intertwined objectives: (i) equipping reinforcement learning with active exploration and deliberate information gathering, such that it regulates state and parameter uncertainties resulting from modeling mismatches and noisy sensory; and (ii) overcoming the computational intractability of stochastic optimal control. We approach both objectives by using reinforcement learning to compute the stochastic optimal control law. On one hand, we avoid the curse of dimensionality prohibiting the direct solution of the stochastic dynamic programming equation. On the other hand, the resulting stochastic optimal control reinforcement learning agent admits caution and probing, that is, optimal online exploration and exploitation. Unlike fixed exploration and exploitation balance, caution and probing are employed automatically by the controller in real-time, even after the learning process is terminated. We conclude the paper with a numerical simulation, illustrating how a Linear Quadratic Regulator with the certainty equivalence assumption may lead to poor performance and filter divergence, while our proposed approach is stabilizing, of an acceptable performance, and computationally convenient.
title Actively Learning Reinforcement Learning: A Stochastic Optimal Control Approach
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2309.10831