Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Denkert, Robert, Pham, Huyên, Warin, Xavier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914776703041536
author Denkert, Robert
Pham, Huyên
Warin, Xavier
author_facet Denkert, Robert
Pham, Huyên
Warin, Xavier
contents We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling applications across various classes of Markovian continuous time control problems, beyond diffusion models, including e.g. regular, impulse and optimal stopping/switching problems. By utilizing change of measure in the control randomisation technique, we derive a new policy gradient representation for these randomised problems, featuring parametrised intensity policies. We further develop actor-critic algorithms specifically designed to address general Markovian stochastic control issues. Our framework is demonstrated through its application to optimal switching problems, with two numerical case studies in the energy sector focusing on real options.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching
Denkert, Robert
Pham, Huyên
Warin, Xavier
Optimization and Control
Machine Learning
93E20 (Primary), 68T07 (Secondary)
We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling applications across various classes of Markovian continuous time control problems, beyond diffusion models, including e.g. regular, impulse and optimal stopping/switching problems. By utilizing change of measure in the control randomisation technique, we derive a new policy gradient representation for these randomised problems, featuring parametrised intensity policies. We further develop actor-critic algorithms specifically designed to address general Markovian stochastic control issues. Our framework is demonstrated through its application to optimal switching problems, with two numerical case studies in the energy sector focusing on real options.
title Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching
topic Optimization and Control
Machine Learning
93E20 (Primary), 68T07 (Secondary)
url https://arxiv.org/abs/2404.17939