Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kayal, Aya, Vakili, Sattar, Toni, Laura, Shiu, Da-shan, Bernacchia, Alberto
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913866108108800
author Kayal, Aya
Vakili, Sattar
Toni, Laura
Shiu, Da-shan
Bernacchia, Alberto
author_facet Kayal, Aya
Vakili, Sattar
Toni, Laura
Shiu, Da-shan
Bernacchia, Alberto
contents Bayesian optimization (BO) with preference-based feedback has recently garnered significant attention due to its emerging applications. We refer to this problem as Bayesian Optimization from Human Feedback (BOHF), which differs from conventional BO by learning the best actions from a reduced feedback model, where only the preference between two actions is revealed to the learner at each time step. The objective is to identify the best action using a limited number of preference queries, typically obtained through costly human feedback. Existing work, which adopts the Bradley-Terry-Luce (BTL) feedback model, provides regret bounds for the performance of several algorithms. In this work, within the same framework we develop tighter performance guarantees. Specifically, we derive regret bounds of $\tilde{\mathcal{O}}(\sqrt{Γ(T)T})$, where $Γ(T)$ represents the maximum information gain$\unicode{x2014}$a kernel-specific complexity term$\unicode{x2014}$and $T$ is the number of queries. Our results significantly improve upon existing bounds. Notably, for common kernels, we show that the order-optimal sample complexities of conventional BO$\unicode{x2014}$achieved with richer feedback models$\unicode{x2014}$are recovered. In other words, the same number of preferential samples as scalar-valued samples is sufficient to find a nearly optimal solution.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23673
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
Kayal, Aya
Vakili, Sattar
Toni, Laura
Shiu, Da-shan
Bernacchia, Alberto
Machine Learning
Bayesian optimization (BO) with preference-based feedback has recently garnered significant attention due to its emerging applications. We refer to this problem as Bayesian Optimization from Human Feedback (BOHF), which differs from conventional BO by learning the best actions from a reduced feedback model, where only the preference between two actions is revealed to the learner at each time step. The objective is to identify the best action using a limited number of preference queries, typically obtained through costly human feedback. Existing work, which adopts the Bradley-Terry-Luce (BTL) feedback model, provides regret bounds for the performance of several algorithms. In this work, within the same framework we develop tighter performance guarantees. Specifically, we derive regret bounds of $\tilde{\mathcal{O}}(\sqrt{Γ(T)T})$, where $Γ(T)$ represents the maximum information gain$\unicode{x2014}$a kernel-specific complexity term$\unicode{x2014}$and $T$ is the number of queries. Our results significantly improve upon existing bounds. Notably, for common kernels, we show that the order-optimal sample complexities of conventional BO$\unicode{x2014}$achieved with richer feedback models$\unicode{x2014}$are recovered. In other words, the same number of preferential samples as scalar-valued samples is sufficient to find a nearly optimal solution.
title Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
topic Machine Learning
url https://arxiv.org/abs/2505.23673