Assessing the impact of dimensionality reduction on clustering performance -- a systematic study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Assani-Amate, Ousmane, Bakhtyari, Mohammadreza, Roy, Émilie, Makarenkov, Vladimir
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910209686896640
author Assani-Amate, Ousmane
Bakhtyari, Mohammadreza
Roy, Émilie
Makarenkov, Vladimir
author_facet Assani-Amate, Ousmane
Bakhtyari, Mohammadreza
Roy, Émilie
Makarenkov, Vladimir
contents Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22099
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Assessing the impact of dimensionality reduction on clustering performance -- a systematic study
Assani-Amate, Ousmane
Bakhtyari, Mohammadreza
Roy, Émilie
Makarenkov, Vladimir
Machine Learning
Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.
title Assessing the impact of dimensionality reduction on clustering performance -- a systematic study
topic Machine Learning
url https://arxiv.org/abs/2604.22099