A Control Theory inspired Exploration Method for a Linear Bandit driven by a Linear Gaussian Dynamical System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gornet, Jonathan, Mo, Yilin, Sinopoli, Bruno
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909820368453632
author Gornet, Jonathan
Mo, Yilin
Sinopoli, Bruno
author_facet Gornet, Jonathan
Mo, Yilin
Sinopoli, Bruno
contents The paper introduces a linear bandit environment where the reward is the output of a known Linear Gaussian Dynamical System (LGDS). In this environment, we address the fundamental challenge of balancing exploration -- gathering information about the environment -- and exploitation -- selecting to the action with the highest predicted reward. We propose two algorithms, Kalman filter Upper Confidence Bound (Kalman-UCB) and Information filter Directed Exploration Action-selection (IDEA). Kalman-UCB uses the principle of optimism in the face of uncertainty. IDEA selects actions that maximize the combination of the predicted reward and a term that quantifies how much an action minimizes the error of the Kalman filter state prediction, which depends on the LGDS property called observability. IDEA is motivated by applications such as hyperparameter optimization in machine learning. A major problem encountered in hyperparameter optimization is the large action spaces, which hinder the performance of methods inspired by principle of optimism in the face of uncertainty as they need to explore each action to lower reward prediction uncertainty. To predict if either Kalman-UCB or IDEA will perform better, a metric based on the LGDS properties is provided. This metric is validated with numerical results across a variety of randomly generated environments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01364
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Control Theory inspired Exploration Method for a Linear Bandit driven by a Linear Gaussian Dynamical System
Gornet, Jonathan
Mo, Yilin
Sinopoli, Bruno
Systems and Control
Signal Processing
The paper introduces a linear bandit environment where the reward is the output of a known Linear Gaussian Dynamical System (LGDS). In this environment, we address the fundamental challenge of balancing exploration -- gathering information about the environment -- and exploitation -- selecting to the action with the highest predicted reward. We propose two algorithms, Kalman filter Upper Confidence Bound (Kalman-UCB) and Information filter Directed Exploration Action-selection (IDEA). Kalman-UCB uses the principle of optimism in the face of uncertainty. IDEA selects actions that maximize the combination of the predicted reward and a term that quantifies how much an action minimizes the error of the Kalman filter state prediction, which depends on the LGDS property called observability. IDEA is motivated by applications such as hyperparameter optimization in machine learning. A major problem encountered in hyperparameter optimization is the large action spaces, which hinder the performance of methods inspired by principle of optimism in the face of uncertainty as they need to explore each action to lower reward prediction uncertainty. To predict if either Kalman-UCB or IDEA will perform better, a metric based on the LGDS properties is provided. This metric is validated with numerical results across a variety of randomly generated environments.
title A Control Theory inspired Exploration Method for a Linear Bandit driven by a Linear Gaussian Dynamical System
topic Systems and Control
Signal Processing
url https://arxiv.org/abs/2510.01364