CDC: A Simple Framework for Complex Data Clustering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Zhao, Xie, Xuanting, Li, Bingheng, Pan, Erlin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909451870535680
author Kang, Zhao
Xie, Xuanting
Li, Bingheng
Pan, Erlin
author_facet Kang, Zhao
Xie, Xuanting
Li, Bingheng
Pan, Erlin
contents In today's data-driven digital era, the amount as well as complexity, such as multi-view, non-Euclidean, and multi-relational, of the collected data are growing exponentially or even faster. Clustering, which unsupervisely extracts valid knowledge from data, is extremely useful in practice. However, existing methods are independently developed to handle one particular challenge at the expense of the others. In this work, we propose a simple but effective framework for complex data clustering (CDC) that can efficiently process different types of data with linear complexity. We first utilize graph filtering to fuse geometry structure and attribute information. We then reduce the complexity with high-quality anchors that are adaptively learned via a novel similarity-preserving regularizer. We illustrate the cluster-ability of our proposed method theoretically and experimentally. In particular, we deploy CDC to graph data of size 111M.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03670
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CDC: A Simple Framework for Complex Data Clustering
Kang, Zhao
Xie, Xuanting
Li, Bingheng
Pan, Erlin
Machine Learning
In today's data-driven digital era, the amount as well as complexity, such as multi-view, non-Euclidean, and multi-relational, of the collected data are growing exponentially or even faster. Clustering, which unsupervisely extracts valid knowledge from data, is extremely useful in practice. However, existing methods are independently developed to handle one particular challenge at the expense of the others. In this work, we propose a simple but effective framework for complex data clustering (CDC) that can efficiently process different types of data with linear complexity. We first utilize graph filtering to fuse geometry structure and attribute information. We then reduce the complexity with high-quality anchors that are adaptively learned via a novel similarity-preserving regularizer. We illustrate the cluster-ability of our proposed method theoretically and experimentally. In particular, we deploy CDC to graph data of size 111M.
title CDC: A Simple Framework for Complex Data Clustering
topic Machine Learning
url https://arxiv.org/abs/2403.03670