A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lazzaro, Joseph, Buffelli, Davide, Shiu, Da-shan, Vakili, Sattar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917442433843200
author Lazzaro, Joseph
Buffelli, Davide
Shiu, Da-shan
Vakili, Sattar
author_facet Lazzaro, Joseph
Buffelli, Davide
Shiu, Da-shan
Vakili, Sattar
contents Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25025
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
Lazzaro, Joseph
Buffelli, Davide
Shiu, Da-shan
Vakili, Sattar
Machine Learning
Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples.
title A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
topic Machine Learning
url https://arxiv.org/abs/2604.25025