An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sutrakar, Vijay Kumar, Mogre, Nikhil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915295937953792
author Sutrakar, Vijay Kumar
Mogre, Nikhil
author_facet Sutrakar, Vijay Kumar
Mogre, Nikhil
contents In this paper, an improved clustering technique for large textual datasets by leveraging fine-tuned word embeddings is presented. WEClustering technique is used as the base model. WEClustering model is fur-ther improvements incorporating fine-tuning contextual embeddings, advanced dimensionality reduction methods, and optimization of clustering algorithms. Experimental results on benchmark datasets demon-strate significant improvements in clustering metrics such as silhouette score, purity, and adjusted rand index (ARI). An increase of 45% and 67% of median silhouette score is reported for the proposed WE-Clustering_K++ (based on K-means) and WEClustering_A++ (based on Agglomerative models), respec-tively. The proposed technique will help to bridge the gap between semantic understanding and statistical robustness for large-scale text-mining tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets
Sutrakar, Vijay Kumar
Mogre, Nikhil
Machine Learning
Computation and Language
In this paper, an improved clustering technique for large textual datasets by leveraging fine-tuned word embeddings is presented. WEClustering technique is used as the base model. WEClustering model is fur-ther improvements incorporating fine-tuning contextual embeddings, advanced dimensionality reduction methods, and optimization of clustering algorithms. Experimental results on benchmark datasets demon-strate significant improvements in clustering metrics such as silhouette score, purity, and adjusted rand index (ARI). An increase of 45% and 67% of median silhouette score is reported for the proposed WE-Clustering_K++ (based on K-means) and WEClustering_A++ (based on Agglomerative models), respec-tively. The proposed technique will help to bridge the gap between semantic understanding and statistical robustness for large-scale text-mining tasks.
title An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2502.16139