Sample-Efficient Reinforcement Learning of Koopman eNMPC

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mayfrank, Daniel, Velioglu, Mehmet, Mitsos, Alexander, Dahmen, Manuel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916734469931008
author Mayfrank, Daniel
Velioglu, Mehmet
Mitsos, Alexander
Dahmen, Manuel
author_facet Mayfrank, Daniel
Velioglu, Mehmet
Mitsos, Alexander
Dahmen, Manuel
contents Reinforcement learning (RL) can be used to tune data-driven (economic) nonlinear model predictive controllers ((e)NMPCs) for optimal performance in a specific control task by optimizing the dynamic model or parameters in the policy's objective function or constraints, such as state bounds. However, the sample efficiency of RL is crucial, and to improve it, we combine a model-based RL algorithm with our published method that turns Koopman (e)NMPCs into automatically differentiable policies. We apply our approach to an eNMPC case study of a continuous stirred-tank reactor (CSTR) model from the literature. The approach outperforms benchmark methods, i.e., data-driven eNMPCs using models based on system identification without further RL tuning of the resulting policy, and neural network controllers trained with model-based RL, by achieving superior control performance and higher sample efficiency. Furthermore, utilizing partial prior knowledge about the system dynamics via physics-informed learning further increases sample efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample-Efficient Reinforcement Learning of Koopman eNMPC
Mayfrank, Daniel
Velioglu, Mehmet
Mitsos, Alexander
Dahmen, Manuel
Machine Learning
Optimization and Control
Reinforcement learning (RL) can be used to tune data-driven (economic) nonlinear model predictive controllers ((e)NMPCs) for optimal performance in a specific control task by optimizing the dynamic model or parameters in the policy's objective function or constraints, such as state bounds. However, the sample efficiency of RL is crucial, and to improve it, we combine a model-based RL algorithm with our published method that turns Koopman (e)NMPCs into automatically differentiable policies. We apply our approach to an eNMPC case study of a continuous stirred-tank reactor (CSTR) model from the literature. The approach outperforms benchmark methods, i.e., data-driven eNMPCs using models based on system identification without further RL tuning of the resulting policy, and neural network controllers trained with model-based RL, by achieving superior control performance and higher sample efficiency. Furthermore, utilizing partial prior knowledge about the system dynamics via physics-informed learning further increases sample efficiency.
title Sample-Efficient Reinforcement Learning of Koopman eNMPC
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2503.18787