Saved in:
Bibliographic Details
Main Authors: Galis, Fabian, Onchis, Darian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.11706
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929761352155136
author Galis, Fabian
Onchis, Darian
author_facet Galis, Fabian
Onchis, Darian
contents In the context of unsupervised learning, effective clustering plays a vital role in revealing patterns and insights from unlabeled data. However, the success of clustering algorithms often depends on the relevance and contribution of features, which can differ between various datasets. This paper explores feature weighting for clustering and presents new weighting strategies, including methods based on SHAP (SHapley Additive exPlanations), a technique commonly used for providing explainability in various supervised machine learning tasks. By taking advantage of SHAP values in a way other than just to gain explainability, we use them to weight features and ultimately improve the clustering process itself in unsupervised scenarios. Our empirical evaluations across five benchmark datasets and clustering methods demonstrate that feature weighting based on SHAP can enhance unsupervised clustering quality, achieving up to a 22.69\% improvement over other weighting methods (from 0.586 to 0.719 in terms of the Adjusted Rand Index). Additionally, these situations where the weighted data boosts the results are highlighted and thoroughly explored, offering insight for practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Refining Filter Global Feature Weighting for Fully-Unsupervised Clustering
Galis, Fabian
Onchis, Darian
Machine Learning
Artificial Intelligence
In the context of unsupervised learning, effective clustering plays a vital role in revealing patterns and insights from unlabeled data. However, the success of clustering algorithms often depends on the relevance and contribution of features, which can differ between various datasets. This paper explores feature weighting for clustering and presents new weighting strategies, including methods based on SHAP (SHapley Additive exPlanations), a technique commonly used for providing explainability in various supervised machine learning tasks. By taking advantage of SHAP values in a way other than just to gain explainability, we use them to weight features and ultimately improve the clustering process itself in unsupervised scenarios. Our empirical evaluations across five benchmark datasets and clustering methods demonstrate that feature weighting based on SHAP can enhance unsupervised clustering quality, achieving up to a 22.69\% improvement over other weighting methods (from 0.586 to 0.719 in terms of the Adjusted Rand Index). Additionally, these situations where the weighted data boosts the results are highlighted and thoroughly explored, offering insight for practical applications.
title Refining Filter Global Feature Weighting for Fully-Unsupervised Clustering
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.11706