Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bertsimas, Dimitris, Kim, Cheol Woo, Niño-Mora, José
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917467117322240
author Bertsimas, Dimitris
Kim, Cheol Woo
Niño-Mora, José
author_facet Bertsimas, Dimitris
Kim, Cheol Woo
Niño-Mora, José
contents We present a novel machine learning framework for the optimal control of fluid restless multi-armed bandit problems (FRMABPs) with state equations that are either affine or quadratic in the state variables. By establishing fundamental properties of FRMABPs, we develop an efficient numerical algorithm that generates a comprehensive training set by solving multiple instances with diverse initial states. We further enhance this training set by applying a nonlinear transformation to the feature vectors, leveraging structural properties of FRMABPs. A time-dependent state feedback policy is then learned using Optimal Classification Trees with hyperplane splits (OCT-H). We test our approach on machine maintenance, epidemic control, and fisheries control problems, demonstrating that our method yields high-quality state feedback policies. Furthermore, once a policy is learned, it achieves a speed-up of up to 26 million times compared to the direct numerical algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach
Bertsimas, Dimitris
Kim, Cheol Woo
Niño-Mora, José
Machine Learning
We present a novel machine learning framework for the optimal control of fluid restless multi-armed bandit problems (FRMABPs) with state equations that are either affine or quadratic in the state variables. By establishing fundamental properties of FRMABPs, we develop an efficient numerical algorithm that generates a comprehensive training set by solving multiple instances with diverse initial states. We further enhance this training set by applying a nonlinear transformation to the feature vectors, leveraging structural properties of FRMABPs. A time-dependent state feedback policy is then learned using Optimal Classification Trees with hyperplane splits (OCT-H). We test our approach on machine maintenance, epidemic control, and fisheries control problems, demonstrating that our method yields high-quality state feedback policies. Furthermore, once a policy is learned, it achieves a speed-up of up to 26 million times compared to the direct numerical algorithm.
title Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach
topic Machine Learning
url https://arxiv.org/abs/2502.03725