Fixed-Confidence Best Arm Identification with Decreasing Variance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roychowdhury, Tamojeet, Reddy, Kota Srinivas, Jagannathan, Krishna P, Moharir, Sharayu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929709782138880
author Roychowdhury, Tamojeet
Reddy, Kota Srinivas
Jagannathan, Krishna P
Moharir, Sharayu
author_facet Roychowdhury, Tamojeet
Reddy, Kota Srinivas
Jagannathan, Krishna P
Moharir, Sharayu
contents We focus on the problem of best-arm identification in a stochastic multi-arm bandit with temporally decreasing variances for the arms' rewards. We model arm rewards as Gaussian random variables with fixed means and variances that decrease with time. The cost incurred by the learner is modeled as a weighted sum of the time needed by the learner to identify the best arm, and the number of samples of arms collected by the learner before termination. Under this cost function, there is an incentive for the learner to not sample arms in all rounds, especially in the initial rounds. On the other hand, not sampling increases the termination time of the learner, which also increases cost. This trade-off necessitates new sampling strategies. We propose two policies. The first policy has an initial wait period with no sampling followed by continuous sampling. The second policy samples periodically and uses a weighted average of the rewards observed to identify the best arm. We provide analytical guarantees on the performance of both policies and supplement our theoretical results with simulations which show that our polices outperform the state-of-the-art policies for the classical best arm identification problem.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07199
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fixed-Confidence Best Arm Identification with Decreasing Variance
Roychowdhury, Tamojeet
Reddy, Kota Srinivas
Jagannathan, Krishna P
Moharir, Sharayu
Machine Learning
Information Theory
Statistics Theory
We focus on the problem of best-arm identification in a stochastic multi-arm bandit with temporally decreasing variances for the arms' rewards. We model arm rewards as Gaussian random variables with fixed means and variances that decrease with time. The cost incurred by the learner is modeled as a weighted sum of the time needed by the learner to identify the best arm, and the number of samples of arms collected by the learner before termination. Under this cost function, there is an incentive for the learner to not sample arms in all rounds, especially in the initial rounds. On the other hand, not sampling increases the termination time of the learner, which also increases cost. This trade-off necessitates new sampling strategies. We propose two policies. The first policy has an initial wait period with no sampling followed by continuous sampling. The second policy samples periodically and uses a weighted average of the rewards observed to identify the best arm. We provide analytical guarantees on the performance of both policies and supplement our theoretical results with simulations which show that our polices outperform the state-of-the-art policies for the classical best arm identification problem.
title Fixed-Confidence Best Arm Identification with Decreasing Variance
topic Machine Learning
Information Theory
Statistics Theory
url https://arxiv.org/abs/2502.07199