Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Brooks, Marc, Durham, Gabriel, Hong, Kihyuk, Tewari, Ambuj
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913852056141824
author Brooks, Marc
Durham, Gabriel
Hong, Kihyuk
Tewari, Ambuj
author_facet Brooks, Marc
Durham, Gabriel
Hong, Kihyuk
Tewari, Ambuj
contents Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using bandit formulations, the integration of GenAI introduces new structure into otherwise classical sequential learning problems. In GenAI-powered interventions, the agent selects a query, but the environment experiences a stochastic response drawn from the generative model. Standard bandit methods do not explicitly account for this structure, where actions influence rewards only through stochastic, observed treatments. We introduce generator-mediated bandit-Thompson sampling (GAMBITTS), a bandit approach designed for this action/treatment split, using mobile health interventions with large language model-generated text as a motivating case study. GAMBITTS explicitly models both the treatment and reward generation processes, using information in the delivered treatment to accelerate policy learning relative to standard methods. We establish regret bounds for GAMBITTS by decomposing sources of uncertainty in treatment and reward, identifying conditions where it achieves stronger guarantees than standard bandit approaches. In simulation studies, GAMBITTS consistently outperforms conventional algorithms by leveraging observed treatments to more accurately estimate expected rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
Brooks, Marc
Durham, Gabriel
Hong, Kihyuk
Tewari, Ambuj
Machine Learning
Methodology
Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using bandit formulations, the integration of GenAI introduces new structure into otherwise classical sequential learning problems. In GenAI-powered interventions, the agent selects a query, but the environment experiences a stochastic response drawn from the generative model. Standard bandit methods do not explicitly account for this structure, where actions influence rewards only through stochastic, observed treatments. We introduce generator-mediated bandit-Thompson sampling (GAMBITTS), a bandit approach designed for this action/treatment split, using mobile health interventions with large language model-generated text as a motivating case study. GAMBITTS explicitly models both the treatment and reward generation processes, using information in the delivered treatment to accelerate policy learning relative to standard methods. We establish regret bounds for GAMBITTS by decomposing sources of uncertainty in treatment and reward, identifying conditions where it achieves stronger guarantees than standard bandit approaches. In simulation studies, GAMBITTS consistently outperforms conventional algorithms by leveraging observed treatments to more accurately estimate expected rewards.
title Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
topic Machine Learning
Methodology
url https://arxiv.org/abs/2505.16311