Invariance-Based Dynamic Regret Minimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lazzaretto, Margherita, Peters, Jonas, Pfister, Niklas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917312699826176
author Lazzaretto, Margherita
Peters, Jonas
Pfister, Niklas
author_facet Lazzaretto, Margherita
Peters, Jonas
Pfister, Niklas
contents We consider stochastic non-stationary linear bandits where the linear parameter connecting contexts to the reward changes over time. Existing algorithms in this setting localize the policy by gradually discarding or down-weighting past data, effectively shrinking the time horizon over which learning can occur. However, in many settings historical data may still carry partial information about the reward model. We propose to leverage such data while adapting to changes, by assuming the reward model decomposes into stationary and non-stationary components. Based on this assumption, we introduce ISD-linUCB, an algorithm that uses past data to learn invariances in the reward model and subsequently exploits them to improve online performance. We show both theoretically and empirically that leveraging invariance reduces the problem dimensionality, yielding significant regret improvements in fast-changing environments when sufficient historical data is available.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03843
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Invariance-Based Dynamic Regret Minimization
Lazzaretto, Margherita
Peters, Jonas
Pfister, Niklas
Machine Learning
We consider stochastic non-stationary linear bandits where the linear parameter connecting contexts to the reward changes over time. Existing algorithms in this setting localize the policy by gradually discarding or down-weighting past data, effectively shrinking the time horizon over which learning can occur. However, in many settings historical data may still carry partial information about the reward model. We propose to leverage such data while adapting to changes, by assuming the reward model decomposes into stationary and non-stationary components. Based on this assumption, we introduce ISD-linUCB, an algorithm that uses past data to learn invariances in the reward model and subsequently exploits them to improve online performance. We show both theoretically and empirically that leveraging invariance reduces the problem dimensionality, yielding significant regret improvements in fast-changing environments when sufficient historical data is available.
title Invariance-Based Dynamic Regret Minimization
topic Machine Learning
url https://arxiv.org/abs/2603.03843