Saved in:
Bibliographic Details
Main Authors: Xie, Hongrui, Cao, Junyu, Xu, Kan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.24231
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917299873644544
author Xie, Hongrui
Cao, Junyu
Xu, Kan
author_facet Xie, Hongrui
Cao, Junyu
Xu, Kan
contents In this paper, we provide the first investigation into adaptive combinatorial experimental design, focusing on the trade-off between regret minimization and statistical power in combinatorial multi-armed bandits (CMAB). While minimizing regret requires repeated exploitation of high-reward arms, accurate inference on reward gaps requires sufficient exploration of suboptimal actions. We formalize this trade-off through the concept of Pareto optimality and establish equivalent conditions for Pareto-efficient learning in CMAB. We consider two relevant cases under different information structures, i.e., full-bandit feedback and semi-bandit feedback, and propose two algorithms MixCombKL and MixCombUCB respectively for these two cases. We provide theoretical guarantees showing that both algorithms are Pareto optimal, achieving finite-time guarantees on both regret and estimation error of arm gaps. Our results further reveal that richer feedback significantly tightens the attainable Pareto frontier, with the primary gains arising from improved estimation accuracy under our proposed methods. Taken together, these findings establish a principled framework for adaptive combinatorial experimentation in multi-objective decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2602_24231
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference
Xie, Hongrui
Cao, Junyu
Xu, Kan
Machine Learning
In this paper, we provide the first investigation into adaptive combinatorial experimental design, focusing on the trade-off between regret minimization and statistical power in combinatorial multi-armed bandits (CMAB). While minimizing regret requires repeated exploitation of high-reward arms, accurate inference on reward gaps requires sufficient exploration of suboptimal actions. We formalize this trade-off through the concept of Pareto optimality and establish equivalent conditions for Pareto-efficient learning in CMAB. We consider two relevant cases under different information structures, i.e., full-bandit feedback and semi-bandit feedback, and propose two algorithms MixCombKL and MixCombUCB respectively for these two cases. We provide theoretical guarantees showing that both algorithms are Pareto optimal, achieving finite-time guarantees on both regret and estimation error of arm gaps. Our results further reveal that richer feedback significantly tightens the attainable Pareto frontier, with the primary gains arising from improved estimation accuracy under our proposed methods. Taken together, these findings establish a principled framework for adaptive combinatorial experimentation in multi-objective decision-making.
title Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference
topic Machine Learning
url https://arxiv.org/abs/2602.24231