Offline Reinforcement Learning via Inverse Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dimanidis, Ioannis, Ok, Tolga, Esfahani, Peyman Mohajerin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912971502911488
author Dimanidis, Ioannis
Ok, Tolga
Esfahani, Peyman Mohajerin
author_facet Dimanidis, Ioannis
Ok, Tolga
Esfahani, Peyman Mohajerin
contents Inspired by the recent successes of Inverse Optimization (IO) across various application domains, we propose a novel offline Reinforcement Learning (ORL) algorithm for continuous state and action spaces, leveraging the convex loss function called ``sub-optimality loss'' from the IO literature. To mitigate the distribution shift commonly observed in ORL problems, we further employ a robust and non-causal Model Predictive Control (MPC) expert steering a nominal model of the dynamics using in-hindsight information stemming from the model mismatch. Unlike the existing literature, our robust MPC expert enjoys an exact and tractable convex reformulation. In the second part of this study, we show that the IO hypothesis class, trained by the proposed convex loss function, enjoys ample expressiveness and {reliably recovers teacher behavior in MuJoCo benchmarks. The method achieves competitive results compared to widely-used baselines in sample-constrained settings, despite using} orders of magnitude fewer parameters. To facilitate the reproducibility of our results, we provide an open-source package implementing the proposed algorithms and the experiments. The code is available at https://github.com/TolgaOk/offlineRLviaIO.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20030
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline Reinforcement Learning via Inverse Optimization
Dimanidis, Ioannis
Ok, Tolga
Esfahani, Peyman Mohajerin
Machine Learning
Systems and Control
Optimization and Control
Inspired by the recent successes of Inverse Optimization (IO) across various application domains, we propose a novel offline Reinforcement Learning (ORL) algorithm for continuous state and action spaces, leveraging the convex loss function called ``sub-optimality loss'' from the IO literature. To mitigate the distribution shift commonly observed in ORL problems, we further employ a robust and non-causal Model Predictive Control (MPC) expert steering a nominal model of the dynamics using in-hindsight information stemming from the model mismatch. Unlike the existing literature, our robust MPC expert enjoys an exact and tractable convex reformulation. In the second part of this study, we show that the IO hypothesis class, trained by the proposed convex loss function, enjoys ample expressiveness and {reliably recovers teacher behavior in MuJoCo benchmarks. The method achieves competitive results compared to widely-used baselines in sample-constrained settings, despite using} orders of magnitude fewer parameters. To facilitate the reproducibility of our results, we provide an open-source package implementing the proposed algorithms and the experiments. The code is available at https://github.com/TolgaOk/offlineRLviaIO.
title Offline Reinforcement Learning via Inverse Optimization
topic Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2502.20030