Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mussabayev, Ravil, Mussabayev, Rustam
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910451962478592
author Mussabayev, Ravil
Mussabayev, Rustam
author_facet Mussabayev, Ravil
Mussabayev, Rustam
contents This paper presents a comparative analysis of different optimization techniques for the K-means algorithm in the context of big data. K-means is a widely used clustering algorithm, but it can suffer from scalability issues when dealing with large datasets. The paper explores different approaches to overcome these issues, including parallelization, approximation, and sampling methods. The authors evaluate the performance of various clustering techniques on a large number of benchmark datasets, comparing them according to the dominance criterion provided by the "less is more" approach (LIMA), i.e., simultaneously along the dimensions of speed, clustering quality, and simplicity. The results show that different techniques are more suitable for different types of datasets and provide insights into the trade-offs between speed and accuracy in K-means clustering for big data. Overall, the paper offers a comprehensive guide for practitioners and researchers on how to optimize K-means for big data applications.
format Preprint
id arxiv_https___arxiv_org_abs_2310_09819
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
Mussabayev, Ravil
Mussabayev, Rustam
Machine Learning
Artificial Intelligence
Optimization and Control
This paper presents a comparative analysis of different optimization techniques for the K-means algorithm in the context of big data. K-means is a widely used clustering algorithm, but it can suffer from scalability issues when dealing with large datasets. The paper explores different approaches to overcome these issues, including parallelization, approximation, and sampling methods. The authors evaluate the performance of various clustering techniques on a large number of benchmark datasets, comparing them according to the dominance criterion provided by the "less is more" approach (LIMA), i.e., simultaneously along the dimensions of speed, clustering quality, and simplicity. The results show that different techniques are more suitable for different types of datasets and provide insights into the trade-offs between speed and accuracy in K-means clustering for big data. Overall, the paper offers a comprehensive guide for practitioners and researchers on how to optimize K-means for big data applications.
title Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2310.09819