Multi-Fidelity Hybrid Reinforcement Learning via Information Gain Maximization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sifaou, Houssem, Simeone, Osvaldo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912592496164864
author Sifaou, Houssem
Simeone, Osvaldo
author_facet Sifaou, Houssem
Simeone, Osvaldo
contents Optimizing a reinforcement learning (RL) policy typically requires extensive interactions with a high-fidelity simulator of the environment, which are often costly or impractical. Offline RL addresses this problem by allowing training from pre-collected data, but its effectiveness is strongly constrained by the size and quality of the dataset. Hybrid offline-online RL leverages both offline data and interactions with a single simulator of the environment. In many real-world scenarios, however, multiple simulators with varying levels of fidelity and computational cost are available. In this work, we study multi-fidelity hybrid RL for policy optimization under a fixed cost budget. We introduce multi-fidelity hybrid RL via information gain maximization (MF-HRL-IGM), a hybrid offline-online RL algorithm that implements fidelity selection based on information gain maximization through a bootstrapping approach. Theoretical analysis establishes the no-regret property of MF-HRL-IGM, while empirical evaluations demonstrate its superior performance compared to existing benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Fidelity Hybrid Reinforcement Learning via Information Gain Maximization
Sifaou, Houssem
Simeone, Osvaldo
Machine Learning
Signal Processing
Optimizing a reinforcement learning (RL) policy typically requires extensive interactions with a high-fidelity simulator of the environment, which are often costly or impractical. Offline RL addresses this problem by allowing training from pre-collected data, but its effectiveness is strongly constrained by the size and quality of the dataset. Hybrid offline-online RL leverages both offline data and interactions with a single simulator of the environment. In many real-world scenarios, however, multiple simulators with varying levels of fidelity and computational cost are available. In this work, we study multi-fidelity hybrid RL for policy optimization under a fixed cost budget. We introduce multi-fidelity hybrid RL via information gain maximization (MF-HRL-IGM), a hybrid offline-online RL algorithm that implements fidelity selection based on information gain maximization through a bootstrapping approach. Theoretical analysis establishes the no-regret property of MF-HRL-IGM, while empirical evaluations demonstrate its superior performance compared to existing benchmarks.
title Multi-Fidelity Hybrid Reinforcement Learning via Information Gain Maximization
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2509.14848