Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Kongchang, Zhang, Tingyu, Chen, Wei, Kong, Fang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911340026658816
author Zhou, Kongchang
Zhang, Tingyu
Chen, Wei
Kong, Fang
author_facet Zhou, Kongchang
Zhang, Tingyu
Chen, Wei
Kong, Fang
contents The problem of combinatorial multi-armed bandits with probabilistically triggered arms (CMAB-T) has been extensively studied. Prior work primarily focuses on either the online setting where an agent learns about the unknown environment through iterative interactions, or the offline setting where a policy is learned solely from logged data. However, each of these paradigms has inherent limitations: online algorithms suffer from high interaction costs and slow adaptation, while offline methods are constrained by dataset quality and lack of exploration capabilities. To address these complementary weaknesses, we propose hybrid CMAB-T, a new framework that integrates offline data with online interaction in a principled manner. Our proposed hybrid CUCB algorithm leverages offline data to guide exploration and accelerate convergence, while strategically incorporating online interactions to mitigate the insufficient coverage or distributional bias of the offline dataset. We provide theoretical guarantees on the algorithm's regret, demonstrating that hybrid CUCB significantly outperforms purely online approaches when high-quality offline data is available, and effectively corrects the bias inherent in offline-only methods when the data is limited or misaligned. Empirical results further demonstrate the consistent advantage of our algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms
Zhou, Kongchang
Zhang, Tingyu
Chen, Wei
Kong, Fang
Machine Learning
The problem of combinatorial multi-armed bandits with probabilistically triggered arms (CMAB-T) has been extensively studied. Prior work primarily focuses on either the online setting where an agent learns about the unknown environment through iterative interactions, or the offline setting where a policy is learned solely from logged data. However, each of these paradigms has inherent limitations: online algorithms suffer from high interaction costs and slow adaptation, while offline methods are constrained by dataset quality and lack of exploration capabilities. To address these complementary weaknesses, we propose hybrid CMAB-T, a new framework that integrates offline data with online interaction in a principled manner. Our proposed hybrid CUCB algorithm leverages offline data to guide exploration and accelerate convergence, while strategically incorporating online interactions to mitigate the insufficient coverage or distributional bias of the offline dataset. We provide theoretical guarantees on the algorithm's regret, demonstrating that hybrid CUCB significantly outperforms purely online approaches when high-quality offline data is available, and effectively corrects the bias inherent in offline-only methods when the data is limited or misaligned. Empirical results further demonstrate the consistent advantage of our algorithm.
title Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms
topic Machine Learning
url https://arxiv.org/abs/2512.21925