Nearly Optimal Best Arm Identification for Semiparametric Bandits

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Kim, Seok-Jin
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918428464381952
author Kim, Seok-Jin
author_facet Kim, Seok-Jin
contents We study fixed-confidence Best Arm Identification (BAI) in semiparametric bandits, where rewards are linear in arm features plus an unknown additive baseline shift. Unlike linear-bandit BAI, this setting requires orthogonalized regression, and its instance-optimal sample complexity has remained open. For the transductive setting, we establish an attainable instance-dependent lower bound characterized by the corresponding linear-bandit complexity on shifted features. We then propose a computationally efficient phase-elimination algorithm based on a new $XY$-design for orthogonalized regression. Our analysis yields a nearly optimal high-probability sample-complexity upper bound, up to log factors and an additive $d^2$ term, and experiments on synthetic instances and the Jester dataset show clear gains over prior baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03969
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Nearly Optimal Best Arm Identification for Semiparametric Bandits
Kim, Seok-Jin
Machine Learning
Methodology
We study fixed-confidence Best Arm Identification (BAI) in semiparametric bandits, where rewards are linear in arm features plus an unknown additive baseline shift. Unlike linear-bandit BAI, this setting requires orthogonalized regression, and its instance-optimal sample complexity has remained open. For the transductive setting, we establish an attainable instance-dependent lower bound characterized by the corresponding linear-bandit complexity on shifted features. We then propose a computationally efficient phase-elimination algorithm based on a new $XY$-design for orthogonalized regression. Our analysis yields a nearly optimal high-probability sample-complexity upper bound, up to log factors and an additive $d^2$ term, and experiments on synthetic instances and the Jester dataset show clear gains over prior baselines.
title Nearly Optimal Best Arm Identification for Semiparametric Bandits
topic Machine Learning
Methodology
url https://arxiv.org/abs/2604.03969