Incentivized Learning in Principal-Agent Bandit Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Scheid, Antoine, Tiapkin, Daniil, Boursier, Etienne, Capitaine, Aymeric, Mhamdi, El Mahdi El, Moulines, Eric, Jordan, Michael I., Durmus, Alain
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916149011152896
author Scheid, Antoine
Tiapkin, Daniil
Boursier, Etienne
Capitaine, Aymeric
Mhamdi, El Mahdi El
Moulines, Eric
Jordan, Michael I.
Durmus, Alain
author_facet Scheid, Antoine
Tiapkin, Daniil
Boursier, Etienne
Capitaine, Aymeric
Mhamdi, El Mahdi El
Moulines, Eric
Jordan, Michael I.
Durmus, Alain
contents This work considers a repeated principal-agent bandit game, where the principal can only interact with her environment through the agent. The principal and the agent have misaligned objectives and the choice of action is only left to the agent. However, the principal can influence the agent's decisions by offering incentives which add up to his rewards. The principal aims to iteratively learn an incentive policy to maximize her own total utility. This framework extends usual bandit problems and is motivated by several practical applications, such as healthcare or ecological taxation, where traditionally used mechanism design theories often overlook the learning aspect of the problem. We present nearly optimal (with respect to a horizon $T$) learning algorithms for the principal's regret in both multi-armed and linear contextual settings. Finally, we support our theoretical guarantees through numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03811
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Incentivized Learning in Principal-Agent Bandit Games
Scheid, Antoine
Tiapkin, Daniil
Boursier, Etienne
Capitaine, Aymeric
Mhamdi, El Mahdi El
Moulines, Eric
Jordan, Michael I.
Durmus, Alain
Machine Learning
Computer Science and Game Theory
This work considers a repeated principal-agent bandit game, where the principal can only interact with her environment through the agent. The principal and the agent have misaligned objectives and the choice of action is only left to the agent. However, the principal can influence the agent's decisions by offering incentives which add up to his rewards. The principal aims to iteratively learn an incentive policy to maximize her own total utility. This framework extends usual bandit problems and is motivated by several practical applications, such as healthcare or ecological taxation, where traditionally used mechanism design theories often overlook the learning aspect of the problem. We present nearly optimal (with respect to a horizon $T$) learning algorithms for the principal's regret in both multi-armed and linear contextual settings. Finally, we support our theoretical guarantees through numerical experiments.
title Incentivized Learning in Principal-Agent Bandit Games
topic Machine Learning
Computer Science and Game Theory
url https://arxiv.org/abs/2403.03811