Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xinmeng, Xu, Kan, Lee, Donghwan, Hassani, Hamed, Bastani, Hamsa, Dobriban, Edgar
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912153858998272
author Huang, Xinmeng
Xu, Kan
Lee, Donghwan
Hassani, Hamed
Bastani, Hamsa
Dobriban, Edgar
author_facet Huang, Xinmeng
Xu, Kan
Lee, Donghwan
Hassani, Hamed
Bastani, Hamsa
Dobriban, Edgar
contents Large and complex datasets are often collected from several, possibly heterogeneous sources. Multitask learning methods improve efficiency by leveraging commonalities across datasets while accounting for possible differences among them. Here, we study multitask linear regression and contextual bandits under sparse heterogeneity, where the source/task-associated parameters are equal to a global parameter plus a sparse task-specific term. We propose a novel two-stage estimator called MOLAR that leverages this structure by first constructing a covariate-wise weighted median of the task-wise linear regression estimates and then shrinking the task-wise estimates towards the weighted median. Compared to task-wise least squares estimates, MOLAR improves the dependence of the estimation error on the data dimension. Extensions of MOLAR to generalized linear models and constructing confidence intervals are discussed in the paper. We then apply MOLAR to develop methods for sparsely heterogeneous multitask contextual bandits, obtaining improved regret guarantees over single-task bandit methods. We further show that our methods are minimax optimal by providing a number of lower bounds. Finally, we support the efficiency of our methods by performing experiments on both synthetic data and the PISA dataset on student educational outcomes from heterogeneous countries.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06291
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
Huang, Xinmeng
Xu, Kan
Lee, Donghwan
Hassani, Hamed
Bastani, Hamsa
Dobriban, Edgar
Machine Learning
Methodology
Large and complex datasets are often collected from several, possibly heterogeneous sources. Multitask learning methods improve efficiency by leveraging commonalities across datasets while accounting for possible differences among them. Here, we study multitask linear regression and contextual bandits under sparse heterogeneity, where the source/task-associated parameters are equal to a global parameter plus a sparse task-specific term. We propose a novel two-stage estimator called MOLAR that leverages this structure by first constructing a covariate-wise weighted median of the task-wise linear regression estimates and then shrinking the task-wise estimates towards the weighted median. Compared to task-wise least squares estimates, MOLAR improves the dependence of the estimation error on the data dimension. Extensions of MOLAR to generalized linear models and constructing confidence intervals are discussed in the paper. We then apply MOLAR to develop methods for sparsely heterogeneous multitask contextual bandits, obtaining improved regret guarantees over single-task bandit methods. We further show that our methods are minimax optimal by providing a number of lower bounds. Finally, we support the efficiency of our methods by performing experiments on both synthetic data and the PISA dataset on student educational outcomes from heterogeneous countries.
title Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
topic Machine Learning
Methodology
url https://arxiv.org/abs/2306.06291