Transfer Learning for Contextual Multi-armed Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Changxiao, Cai, T. Tony, Li, Hongzhe
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909082451968000
author Cai, Changxiao
Cai, T. Tony
Li, Hongzhe
author_facet Cai, Changxiao
Cai, T. Tony
Li, Hongzhe
contents Motivated by a range of applications, we study in this paper the problem of transfer learning for nonparametric contextual multi-armed bandits under the covariate shift model, where we have data collected on source bandits before the start of the target bandit learning. The minimax rate of convergence for the cumulative regret is established and a novel transfer learning algorithm that attains the minimax regret is proposed. The results quantify the contribution of the data from the source domains for learning in the target domain in the context of nonparametric contextual multi-armed bandits. In view of the general impossibility of adaptation to unknown smoothness, we develop a data-driven algorithm that achieves near-optimal statistical guarantees (up to a logarithmic factor) while automatically adapting to the unknown parameters over a large collection of parameter spaces under an additional self-similarity assumption. A simulation study is carried out to illustrate the benefits of utilizing the data from the auxiliary source domains for learning in the target domain.
format Preprint
id arxiv_https___arxiv_org_abs_2211_12612
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Transfer Learning for Contextual Multi-armed Bandits
Cai, Changxiao
Cai, T. Tony
Li, Hongzhe
Machine Learning
Statistics Theory
Motivated by a range of applications, we study in this paper the problem of transfer learning for nonparametric contextual multi-armed bandits under the covariate shift model, where we have data collected on source bandits before the start of the target bandit learning. The minimax rate of convergence for the cumulative regret is established and a novel transfer learning algorithm that attains the minimax regret is proposed. The results quantify the contribution of the data from the source domains for learning in the target domain in the context of nonparametric contextual multi-armed bandits. In view of the general impossibility of adaptation to unknown smoothness, we develop a data-driven algorithm that achieves near-optimal statistical guarantees (up to a logarithmic factor) while automatically adapting to the unknown parameters over a large collection of parameter spaces under an additional self-similarity assumption. A simulation study is carried out to illustrate the benefits of utilizing the data from the auxiliary source domains for learning in the target domain.
title Transfer Learning for Contextual Multi-armed Bandits
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2211.12612