HR-Bandit: Human-AI Collaborated Linear Recourse Bandit

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cao, Junyu, Gao, Ruijiang, Keyvanshokooh, Esmaeil
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913746614484992
author Cao, Junyu
Gao, Ruijiang
Keyvanshokooh, Esmaeil
author_facet Cao, Junyu
Gao, Ruijiang
Keyvanshokooh, Esmaeil
contents Human doctors frequently recommend actionable recourses that allow patients to modify their conditions to access more effective treatments. Inspired by such healthcare scenarios, we propose the Recourse Linear UCB ($\textsf{RLinUCB}$) algorithm, which optimizes both action selection and feature modifications by balancing exploration and exploitation. We further extend this to the Human-AI Linear Recourse Bandit ($\textsf{HR-Bandit}$), which integrates human expertise to enhance performance. $\textsf{HR-Bandit}$ offers three key guarantees: (i) a warm-start guarantee for improved initial performance, (ii) a human-effort guarantee to minimize required human interactions, and (iii) a robustness guarantee that ensures sublinear regret even when human decisions are suboptimal. Empirical results, including a healthcare case study, validate its superior performance against existing benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14640
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HR-Bandit: Human-AI Collaborated Linear Recourse Bandit
Cao, Junyu
Gao, Ruijiang
Keyvanshokooh, Esmaeil
Machine Learning
Human doctors frequently recommend actionable recourses that allow patients to modify their conditions to access more effective treatments. Inspired by such healthcare scenarios, we propose the Recourse Linear UCB ($\textsf{RLinUCB}$) algorithm, which optimizes both action selection and feature modifications by balancing exploration and exploitation. We further extend this to the Human-AI Linear Recourse Bandit ($\textsf{HR-Bandit}$), which integrates human expertise to enhance performance. $\textsf{HR-Bandit}$ offers three key guarantees: (i) a warm-start guarantee for improved initial performance, (ii) a human-effort guarantee to minimize required human interactions, and (iii) a robustness guarantee that ensures sublinear regret even when human decisions are suboptimal. Empirical results, including a healthcare case study, validate its superior performance against existing benchmarks.
title HR-Bandit: Human-AI Collaborated Linear Recourse Bandit
topic Machine Learning
url https://arxiv.org/abs/2410.14640