Collaborating in Multi-Armed Bandits with Strategic Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Barnea, Idan, Schlisselberg, Ofir, Mansour, Yishay
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910215620788224
author Barnea, Idan
Schlisselberg, Ofir
Mansour, Yishay
author_facet Barnea, Idan
Schlisselberg, Ofir
Mansour, Yishay
contents We study collaborative learning in multi-agent Bayesian bandit problems, where strategic agents collectively solve the same bandit instance. While multiple agents can accelerate learning by sharing information, strategic agents might prefer to free-ride and avoid exploration. We consider a setting with persistent agents that participate in multiple time periods. This is in contrast to most previous works on incentives in multi-agent MAB, which assume short-lived agents, namely each agent has a single decision to make and optimizes their expected reward in that single decision. As in the multi-agent MAB model with incentives, our model does not have monetary transfers, and the only incentives are through information sharing. We propose \texttt{CAOS}, a mechanism that sustains collaboration as a Nash equilibrium while achieving strong regret guarantees. Our results demonstrate that collaborative exploration can be sustained purely through information sharing, achieving performance close to that of fully cooperative systems despite strategic behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13145
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Collaborating in Multi-Armed Bandits with Strategic Agents
Barnea, Idan
Schlisselberg, Ofir
Mansour, Yishay
Machine Learning
We study collaborative learning in multi-agent Bayesian bandit problems, where strategic agents collectively solve the same bandit instance. While multiple agents can accelerate learning by sharing information, strategic agents might prefer to free-ride and avoid exploration. We consider a setting with persistent agents that participate in multiple time periods. This is in contrast to most previous works on incentives in multi-agent MAB, which assume short-lived agents, namely each agent has a single decision to make and optimizes their expected reward in that single decision. As in the multi-agent MAB model with incentives, our model does not have monetary transfers, and the only incentives are through information sharing. We propose \texttt{CAOS}, a mechanism that sustains collaboration as a Nash equilibrium while achieving strong regret guarantees. Our results demonstrate that collaborative exploration can be sustained purely through information sharing, achieving performance close to that of fully cooperative systems despite strategic behavior.
title Collaborating in Multi-Armed Bandits with Strategic Agents
topic Machine Learning
url https://arxiv.org/abs/2605.13145