Regret Bounds for Robust Online Decision Making

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Appel, Alexander, Kosoy, Vanessa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911023894626304
author Appel, Alexander
Kosoy, Vanessa
author_facet Appel, Alexander
Kosoy, Vanessa
contents We propose a framework which generalizes "decision making with structured observations" by allowing robust (i.e. multivalued) models. In this framework, each model associates each decision with a convex set of probability distributions over outcomes. Nature can choose distributions out of this set in an arbitrary (adversarial) manner, that can be nonoblivious and depend on past history. The resulting framework offers much greater generality than classical bandits and reinforcement learning, since the realizability assumption becomes much weaker and more realistic. We then derive a theory of regret bounds for this framework. Although our lower and upper bounds are not tight, they are sufficient to fully characterize power-law learnability. We demonstrate this theory in two special cases: robust linear bandits and tabular robust online reinforcement learning. In both cases, we derive regret bounds that improve state-of-the-art (except that we do not address computational efficiency).
format Preprint
id arxiv_https___arxiv_org_abs_2504_06820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Regret Bounds for Robust Online Decision Making
Appel, Alexander
Kosoy, Vanessa
Machine Learning
68Q32
I.2.6
We propose a framework which generalizes "decision making with structured observations" by allowing robust (i.e. multivalued) models. In this framework, each model associates each decision with a convex set of probability distributions over outcomes. Nature can choose distributions out of this set in an arbitrary (adversarial) manner, that can be nonoblivious and depend on past history. The resulting framework offers much greater generality than classical bandits and reinforcement learning, since the realizability assumption becomes much weaker and more realistic. We then derive a theory of regret bounds for this framework. Although our lower and upper bounds are not tight, they are sufficient to fully characterize power-law learnability. We demonstrate this theory in two special cases: robust linear bandits and tabular robust online reinforcement learning. In both cases, we derive regret bounds that improve state-of-the-art (except that we do not address computational efficiency).
title Regret Bounds for Robust Online Decision Making
topic Machine Learning
68Q32
I.2.6
url https://arxiv.org/abs/2504.06820