On uniqueness of the set of k-means

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cárcamo, Javier, Cuevas, Antonio, Rodríguez, Luis A.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916442834731008
author Cárcamo, Javier
Cuevas, Antonio
Rodríguez, Luis A.
author_facet Cárcamo, Javier
Cuevas, Antonio
Rodríguez, Luis A.
contents We provide necessary and sufficient conditions for the uniqueness of the k-means set of a probability distribution. This uniqueness problem is related to the choice of k: depending on the underlying distribution, some values of this parameter could lead to multiple sets of k-means, which hampers the interpretation of the results and/or the stability of the algorithms. We give a general assessment on consistency of the empirical k-means adapted to the setting of non-uniqueness and determine the asymptotic distribution of the within cluster sum of squares (WCSS). We also provide statistical characterizations of k-means uniqueness in terms of the asymptotic behavior of the empirical WCSS. As a consequence, we derive a bootstrap test for uniqueness of the set of k-means. The results are illustrated with examples of different types of non-uniqueness and we check by simulations the performance of the proposed methodology.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On uniqueness of the set of k-means
Cárcamo, Javier
Cuevas, Antonio
Rodríguez, Luis A.
Statistics Theory
Methodology
Machine Learning
62H30, 62E20 (primary), 62G20 (secondary)
We provide necessary and sufficient conditions for the uniqueness of the k-means set of a probability distribution. This uniqueness problem is related to the choice of k: depending on the underlying distribution, some values of this parameter could lead to multiple sets of k-means, which hampers the interpretation of the results and/or the stability of the algorithms. We give a general assessment on consistency of the empirical k-means adapted to the setting of non-uniqueness and determine the asymptotic distribution of the within cluster sum of squares (WCSS). We also provide statistical characterizations of k-means uniqueness in terms of the asymptotic behavior of the empirical WCSS. As a consequence, we derive a bootstrap test for uniqueness of the set of k-means. The results are illustrated with examples of different types of non-uniqueness and we check by simulations the performance of the proposed methodology.
title On uniqueness of the set of k-means
topic Statistics Theory
Methodology
Machine Learning
62H30, 62E20 (primary), 62G20 (secondary)
url https://arxiv.org/abs/2410.13495