A Pontryagin Perspective on Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eberhard, Onno, Vernade, Claire, Muehlebach, Michael
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915252712505344
author Eberhard, Onno
Vernade, Claire
Muehlebach, Michael
author_facet Eberhard, Onno
Vernade, Claire
Muehlebach, Michael
contents Reinforcement learning has traditionally focused on learning state-dependent policies to solve optimal control problems in a closed-loop fashion. In this work, we introduce the paradigm of open-loop reinforcement learning where a fixed action sequence is learned instead. We present three new algorithms: one robust model-based method and two sample-efficient model-free methods. Rather than basing our algorithms on Bellman's equation from dynamic programming, our work builds on Pontryagin's principle from the theory of open-loop optimal control. We provide convergence guarantees and evaluate all methods empirically on a pendulum swing-up task, as well as on two high-dimensional MuJoCo tasks, significantly outperforming existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18100
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Pontryagin Perspective on Reinforcement Learning
Eberhard, Onno
Vernade, Claire
Muehlebach, Michael
Machine Learning
Optimization and Control
Reinforcement learning has traditionally focused on learning state-dependent policies to solve optimal control problems in a closed-loop fashion. In this work, we introduce the paradigm of open-loop reinforcement learning where a fixed action sequence is learned instead. We present three new algorithms: one robust model-based method and two sample-efficient model-free methods. Rather than basing our algorithms on Bellman's equation from dynamic programming, our work builds on Pontryagin's principle from the theory of open-loop optimal control. We provide convergence guarantees and evaluate all methods empirically on a pendulum swing-up task, as well as on two high-dimensional MuJoCo tasks, significantly outperforming existing baselines.
title A Pontryagin Perspective on Reinforcement Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2405.18100