Flexible Variable Selection for Clustering and Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neal, Mackenzie R., McNicholas, Paul D.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910324960002048
author Neal, Mackenzie R.
McNicholas, Paul D.
author_facet Neal, Mackenzie R.
McNicholas, Paul D.
contents The importance of variable selection for clustering has been recognized for some time, and mixture models are well-established as a statistical approach to clustering. Yet, the literature on variable selection in model-based clustering remains largely rooted in the assumption of Gaussian clusters. Unsurprisingly, variable selection algorithms based on this assumption tend to break down in the presence of cluster skewness. A novel variable selection algorithm is presented that utilizes the Manly transformation mixture model to select variables based on their ability to separate clusters, and is effective even when clusters depart from the Gaussian assumption. The proposed approach, which is implemented within the R package vscc, is compared to existing variable selection methods -- including an existing method that can account for cluster skewness -- using simulated and real datasets
format Preprint
id arxiv_https___arxiv_org_abs_2305_16464
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Flexible Variable Selection for Clustering and Classification
Neal, Mackenzie R.
McNicholas, Paul D.
Methodology
The importance of variable selection for clustering has been recognized for some time, and mixture models are well-established as a statistical approach to clustering. Yet, the literature on variable selection in model-based clustering remains largely rooted in the assumption of Gaussian clusters. Unsurprisingly, variable selection algorithms based on this assumption tend to break down in the presence of cluster skewness. A novel variable selection algorithm is presented that utilizes the Manly transformation mixture model to select variables based on their ability to separate clusters, and is effective even when clusters depart from the Gaussian assumption. The proposed approach, which is implemented within the R package vscc, is compared to existing variable selection methods -- including an existing method that can account for cluster skewness -- using simulated and real datasets
title Flexible Variable Selection for Clustering and Classification
topic Methodology
url https://arxiv.org/abs/2305.16464