Partially Observable Contextual Bandits with Linear Payoffs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Sihan, Bhatt, Sujay, Koppel, Alec, Ganesh, Sumitra
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917779119013888
author Zeng, Sihan
Bhatt, Sujay
Koppel, Alec
Ganesh, Sumitra
author_facet Zeng, Sihan
Bhatt, Sujay
Koppel, Alec
Ganesh, Sumitra
contents The standard contextual bandit framework assumes fully observable and actionable contexts. In this work, we consider a new bandit setting with partially observable, correlated contexts and linear payoffs, motivated by the applications in finance where decision making is based on market information that typically displays temporal correlation and is not fully observed. We make the following contributions marrying ideas from statistical signal processing with bandits: (i) We propose an algorithmic pipeline named EMKF-Bandit, which integrates system identification, filtering, and classic contextual bandit algorithms into an iterative method alternating between latent parameter estimation and decision making. (ii) We analyze EMKF-Bandit when we select Thompson sampling as the bandit algorithm and show that it incurs a sub-linear regret under conditions on filtering. (iii) We conduct numerical simulations that demonstrate the benefits and practical applicability of the proposed pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Partially Observable Contextual Bandits with Linear Payoffs
Zeng, Sihan
Bhatt, Sujay
Koppel, Alec
Ganesh, Sumitra
Machine Learning
The standard contextual bandit framework assumes fully observable and actionable contexts. In this work, we consider a new bandit setting with partially observable, correlated contexts and linear payoffs, motivated by the applications in finance where decision making is based on market information that typically displays temporal correlation and is not fully observed. We make the following contributions marrying ideas from statistical signal processing with bandits: (i) We propose an algorithmic pipeline named EMKF-Bandit, which integrates system identification, filtering, and classic contextual bandit algorithms into an iterative method alternating between latent parameter estimation and decision making. (ii) We analyze EMKF-Bandit when we select Thompson sampling as the bandit algorithm and show that it incurs a sub-linear regret under conditions on filtering. (iii) We conduct numerical simulations that demonstrate the benefits and practical applicability of the proposed pipeline.
title Partially Observable Contextual Bandits with Linear Payoffs
topic Machine Learning
url https://arxiv.org/abs/2409.11521