Binary $k$-Center with Missing Entries: Structure Leads to Tractability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Soheil, Farehe, Simonov, Kirill, Friedrich, Tobias
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929739197841408
author Soheil, Farehe
Simonov, Kirill
Friedrich, Tobias
author_facet Soheil, Farehe
Simonov, Kirill
Friedrich, Tobias
contents $\kC$ clustering is a fundamental classification problem, where the task is to categorize the given collection of entities into $k$ clusters and come up with a representative for each cluster, so that the maximum distance between an entity and its representative is minimized. In this work, we focus on the setting where the entities are represented by binary vectors with missing entries, which model incomplete categorical data. This version of the problem has wide applications, from predictive analytics to bioinformatics. Our main finding is that the problem, which is notoriously hard from the classical complexity viewpoint, becomes tractable as soon as the known entries are sparse and exhibit a certain structure. Formally, we show fixed-parameter tractable algorithms for the parameters vertex cover, fracture number, and treewidth of the row-column graph, which encodes the positions of the known entries of the matrix. Additionally, we tie the complexity of the 1-cluster variant of the problem, which is famous under the name Closest String, to the complexity of solving integer linear programs with few constraints. This implies, in particular, that improving upon the running times of our algorithms would lead to more efficient algorithms for integer linear programming in general.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01445
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Binary $k$-Center with Missing Entries: Structure Leads to Tractability
Soheil, Farehe
Simonov, Kirill
Friedrich, Tobias
Data Structures and Algorithms
$\kC$ clustering is a fundamental classification problem, where the task is to categorize the given collection of entities into $k$ clusters and come up with a representative for each cluster, so that the maximum distance between an entity and its representative is minimized. In this work, we focus on the setting where the entities are represented by binary vectors with missing entries, which model incomplete categorical data. This version of the problem has wide applications, from predictive analytics to bioinformatics. Our main finding is that the problem, which is notoriously hard from the classical complexity viewpoint, becomes tractable as soon as the known entries are sparse and exhibit a certain structure. Formally, we show fixed-parameter tractable algorithms for the parameters vertex cover, fracture number, and treewidth of the row-column graph, which encodes the positions of the known entries of the matrix. Additionally, we tie the complexity of the 1-cluster variant of the problem, which is famous under the name Closest String, to the complexity of solving integer linear programs with few constraints. This implies, in particular, that improving upon the running times of our algorithms would lead to more efficient algorithms for integer linear programming in general.
title Binary $k$-Center with Missing Entries: Structure Leads to Tractability
topic Data Structures and Algorithms
url https://arxiv.org/abs/2503.01445