Multi-Task Representation Learning for Conservative Linear Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jiabin, Moothedath, Shana
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911675341340672
author Lin, Jiabin
Moothedath, Shana
author_facet Lin, Jiabin
Moothedath, Shana
contents This paper presents the Constrained Multi-Task Representation Learning (CMTRL) framework for linear bandits. We consider T linear bandit tasks in a d dimensional space, which share a common low-dimensional representation of dimension r, where r is much smaller than the minimum of d and T. Furthermore, tasks are constrained so that only actions meeting specific safety or performance requirements are allowed, referred to as conservative (safe) bandits. We introduce a novel algorithm, Safe-Alternating projected Gradient Descent and minimization (Safe-AltGDmin), to recover a low-rank feature matrix while satisfying the given constraints. Building on this algorithm, we propose a multi-task representation learning framework for conservative linear bandits and establish theoretical guarantees for its regret and sample complexity bounds. We presented experiments and compared the performance of our algorithm with benchmark algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12176
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-Task Representation Learning for Conservative Linear Bandits
Lin, Jiabin
Moothedath, Shana
Machine Learning
This paper presents the Constrained Multi-Task Representation Learning (CMTRL) framework for linear bandits. We consider T linear bandit tasks in a d dimensional space, which share a common low-dimensional representation of dimension r, where r is much smaller than the minimum of d and T. Furthermore, tasks are constrained so that only actions meeting specific safety or performance requirements are allowed, referred to as conservative (safe) bandits. We introduce a novel algorithm, Safe-Alternating projected Gradient Descent and minimization (Safe-AltGDmin), to recover a low-rank feature matrix while satisfying the given constraints. Building on this algorithm, we propose a multi-task representation learning framework for conservative linear bandits and establish theoretical guarantees for its regret and sample complexity bounds. We presented experiments and compared the performance of our algorithm with benchmark algorithms.
title Multi-Task Representation Learning for Conservative Linear Bandits
topic Machine Learning
url https://arxiv.org/abs/2605.12176