Operator Models for Continuous-Time Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoischen, Nicolas, Bevanda, Petar, Beier, Max, Sosnowski, Stefan, Houska, Boris, Hirche, Sandra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912707057287168
author Hoischen, Nicolas
Bevanda, Petar
Beier, Max
Sosnowski, Stefan
Houska, Boris
Hirche, Sandra
author_facet Hoischen, Nicolas
Bevanda, Petar
Beier, Max
Sosnowski, Stefan
Houska, Boris
Hirche, Sandra
contents Continuous-time stochastic processes underlie many natural and engineered systems. In healthcare, autonomous driving, and industrial control, direct interaction with the environment is often unsafe or impractical, motivating offline reinforcement learning from historical data. However, there is limited statistical understanding of the approximation errors inherent in learning policies from offline datasets. We address this by linking reinforcement learning to the Hamilton-Jacobi-Bellman equation and proposing an operator-theoretic algorithm based on a simple dynamic programming recursion. Specifically, we represent our world model in terms of the infinitesimal generator of controlled diffusion processes learned in a reproducing kernel Hilbert space. By integrating statistical learning methods and operator theory, we establish global convergence of the value function and derive finite-sample guarantees with bounds tied to system properties such as smoothness and stability. Our theoretical and numerical results indicate that operator-based approaches may hold promise in solving offline reinforcement learning using continuous-time optimal control.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Operator Models for Continuous-Time Offline Reinforcement Learning
Hoischen, Nicolas
Bevanda, Petar
Beier, Max
Sosnowski, Stefan
Houska, Boris
Hirche, Sandra
Machine Learning
Systems and Control
Optimization and Control
Continuous-time stochastic processes underlie many natural and engineered systems. In healthcare, autonomous driving, and industrial control, direct interaction with the environment is often unsafe or impractical, motivating offline reinforcement learning from historical data. However, there is limited statistical understanding of the approximation errors inherent in learning policies from offline datasets. We address this by linking reinforcement learning to the Hamilton-Jacobi-Bellman equation and proposing an operator-theoretic algorithm based on a simple dynamic programming recursion. Specifically, we represent our world model in terms of the infinitesimal generator of controlled diffusion processes learned in a reproducing kernel Hilbert space. By integrating statistical learning methods and operator theory, we establish global convergence of the value function and derive finite-sample guarantees with bounds tied to system properties such as smoothness and stability. Our theoretical and numerical results indicate that operator-based approaches may hold promise in solving offline reinforcement learning using continuous-time optimal control.
title Operator Models for Continuous-Time Offline Reinforcement Learning
topic Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2511.10383