Query-Efficient Correlation Clustering with Noisy Oracle

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kuroki, Yuko, Miyauchi, Atsushi, Bonchi, Francesco, Chen, Wei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910680960991232
author Kuroki, Yuko
Miyauchi, Atsushi
Bonchi, Francesco
Chen, Wei
author_facet Kuroki, Yuko
Miyauchi, Atsushi
Bonchi, Francesco
Chen, Wei
contents We study a general clustering setting in which we have $n$ elements to be clustered, and we aim to perform as few queries as possible to an oracle that returns a noisy sample of the weighted similarity between two elements. Our setting encompasses many application domains in which the similarity function is costly to compute and inherently noisy. We introduce two novel formulations of online learning problems rooted in the paradigm of Pure Exploration in Combinatorial Multi-Armed Bandits (PE-CMAB): fixed confidence and fixed budget settings. For both settings, we design algorithms that combine a sampling strategy with a classic approximation algorithm for correlation clustering and study their theoretical guarantees. Our results are the first examples of polynomial-time algorithms that work for the case of PE-CMAB in which the underlying offline optimization problem is NP-hard.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01400
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Query-Efficient Correlation Clustering with Noisy Oracle
Kuroki, Yuko
Miyauchi, Atsushi
Bonchi, Francesco
Chen, Wei
Machine Learning
Data Structures and Algorithms
We study a general clustering setting in which we have $n$ elements to be clustered, and we aim to perform as few queries as possible to an oracle that returns a noisy sample of the weighted similarity between two elements. Our setting encompasses many application domains in which the similarity function is costly to compute and inherently noisy. We introduce two novel formulations of online learning problems rooted in the paradigm of Pure Exploration in Combinatorial Multi-Armed Bandits (PE-CMAB): fixed confidence and fixed budget settings. For both settings, we design algorithms that combine a sampling strategy with a classic approximation algorithm for correlation clustering and study their theoretical guarantees. Our results are the first examples of polynomial-time algorithms that work for the case of PE-CMAB in which the underlying offline optimization problem is NP-hard.
title Query-Efficient Correlation Clustering with Noisy Oracle
topic Machine Learning
Data Structures and Algorithms
url https://arxiv.org/abs/2402.01400