Saved in:
Bibliographic Details
Main Authors: Zhou, Yijia, Gallivan, Kyle A., Barbu, Adrian
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2302.14599
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917802636476416
author Zhou, Yijia
Gallivan, Kyle A.
Barbu, Adrian
author_facet Zhou, Yijia
Gallivan, Kyle A.
Barbu, Adrian
contents Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a provably robust clustering algorithm based on loss minimization that performs well on Gaussian mixture models with outliers. It provides theoretical guarantees that the algorithm obtains high accuracy with high probability under certain assumptions. Moreover, it can also be used as an initialization strategy for $k$-means clustering. Experiments on real-world large-scale datasets demonstrate the effectiveness of the algorithm when clustering a large number of clusters, and a $k$-means algorithm initialized by the algorithm outperforms many of the classic clustering methods in both speed and accuracy, while scaling well to large datasets such as ImageNet.
format Preprint
id arxiv_https___arxiv_org_abs_2302_14599
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers
Zhou, Yijia
Gallivan, Kyle A.
Barbu, Adrian
Machine Learning
Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a provably robust clustering algorithm based on loss minimization that performs well on Gaussian mixture models with outliers. It provides theoretical guarantees that the algorithm obtains high accuracy with high probability under certain assumptions. Moreover, it can also be used as an initialization strategy for $k$-means clustering. Experiments on real-world large-scale datasets demonstrate the effectiveness of the algorithm when clustering a large number of clusters, and a $k$-means algorithm initialized by the algorithm outperforms many of the classic clustering methods in both speed and accuracy, while scaling well to large datasets such as ImageNet.
title Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers
topic Machine Learning
url https://arxiv.org/abs/2302.14599