Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mussabayev, Rustam, Mussabayev, Ravil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916180463190016
author Mussabayev, Rustam
Mussabayev, Ravil
author_facet Mussabayev, Rustam
Mussabayev, Ravil
contents This paper introduces a novel K-means clustering algorithm, an advancement on the conventional Big-means methodology. The proposed method efficiently integrates parallel processing, stochastic sampling, and competitive optimization to create a scalable variant designed for big data applications. It addresses scalability and computation time challenges typically faced with traditional techniques. The algorithm adjusts sample sizes dynamically for each worker during execution, optimizing performance. Data from these sample sizes are continually analyzed, facilitating the identification of the most efficient configuration. By incorporating a competitive element among workers using different sample sizes, efficiency within the Big-means algorithm is further stimulated. In essence, the algorithm balances computational time and clustering quality by employing a stochastic, competitive sampling strategy in a parallel computing setting.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18766
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
Mussabayev, Rustam
Mussabayev, Ravil
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Information Retrieval
This paper introduces a novel K-means clustering algorithm, an advancement on the conventional Big-means methodology. The proposed method efficiently integrates parallel processing, stochastic sampling, and competitive optimization to create a scalable variant designed for big data applications. It addresses scalability and computation time challenges typically faced with traditional techniques. The algorithm adjusts sample sizes dynamically for each worker during execution, optimizing performance. Data from these sample sizes are continually analyzed, facilitating the identification of the most efficient configuration. By incorporating a competitive element among workers using different sample sizes, efficiency within the Big-means algorithm is further stimulated. In essence, the algorithm balances computational time and clustering quality by employing a stochastic, competitive sampling strategy in a parallel computing setting.
title Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Information Retrieval
url https://arxiv.org/abs/2403.18766