Mixed-type Distance Shrinkage and Selection for Clustering via Kernel Metric Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghashti, Jesse S., Thompson, John R. J.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914969664094208
author Ghashti, Jesse S.
Thompson, John R. J.
author_facet Ghashti, Jesse S.
Thompson, John R. J.
contents Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity. While there exist numerous distance-based measures for data with pure numerical attributes and several ordered and unordered categorical metrics, an efficient and accurate distance for mixed-type data that utilizes the continuous and discrete properties simulatenously is an open problem. Many metrics convert numerical attributes to categorical ones or vice versa. They handle the data points as a single attribute type or calculate a distance between each attribute separately and add them up. We propose a metric called KDSUM that uses mixed kernels to measure dissimilarity, with cross-validated optimal bandwidth selection. We demonstrate that KDSUM is a shrinkage method from existing mixed-type metrics to a uniform dissimilarity metric, and improves clustering accuracy when utilized in existing distance-based clustering algorithms on simulated and real-world datasets containing continuous-only, categorical-only, and mixed-type data.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01890
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mixed-type Distance Shrinkage and Selection for Clustering via Kernel Metric Learning
Ghashti, Jesse S.
Thompson, John R. J.
Machine Learning
Computation
Methodology
Other Statistics
62G07, 65D10
I.5.3; G.3; I.6.6
Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity. While there exist numerous distance-based measures for data with pure numerical attributes and several ordered and unordered categorical metrics, an efficient and accurate distance for mixed-type data that utilizes the continuous and discrete properties simulatenously is an open problem. Many metrics convert numerical attributes to categorical ones or vice versa. They handle the data points as a single attribute type or calculate a distance between each attribute separately and add them up. We propose a metric called KDSUM that uses mixed kernels to measure dissimilarity, with cross-validated optimal bandwidth selection. We demonstrate that KDSUM is a shrinkage method from existing mixed-type metrics to a uniform dissimilarity metric, and improves clustering accuracy when utilized in existing distance-based clustering algorithms on simulated and real-world datasets containing continuous-only, categorical-only, and mixed-type data.
title Mixed-type Distance Shrinkage and Selection for Clustering via Kernel Metric Learning
topic Machine Learning
Computation
Methodology
Other Statistics
62G07, 65D10
I.5.3; G.3; I.6.6
url https://arxiv.org/abs/2306.01890