Mass Distribution versus Density Distribution in the Context of Clustering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ting, Kai Ming, Zhu, Ye, Zhang, Hang, Liang, Tianrun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911393191559168
author Ting, Kai Ming
Zhu, Ye
Zhang, Hang
Liang, Tianrun
author_facet Ting, Kai Ming
Zhu, Ye
Zhang, Hang
Liang, Tianrun
contents This paper investigates two fundamental descriptors of data, i.e., density distribution versus mass distribution, in the context of clustering. Density distribution has been the de facto descriptor of data distribution since the introduction of statistics. We show that density distribution has its fundamental limitation -- high-density bias, irrespective of the algorithms used to perform clustering. Existing density-based clustering algorithms have employed different algorithmic means to counter the effect of the high-density bias with some success, but the fundamental limitation of using density distribution remains an obstacle to discovering clusters of arbitrary shapes, sizes and densities. Using the mass distribution as a better foundation, we propose a new algorithm which maximizes the total mass of all clusters, called mass-maximization clustering (MMC). The algorithm can be easily changed to maximize the total density of all clusters in order to examine the fundamental limitation of using density distribution versus mass distribution. The key advantage of the MMC over the density-maximization clustering is that the maximization is conducted without a bias towards dense clusters.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10759
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mass Distribution versus Density Distribution in the Context of Clustering
Ting, Kai Ming
Zhu, Ye
Zhang, Hang
Liang, Tianrun
Machine Learning
This paper investigates two fundamental descriptors of data, i.e., density distribution versus mass distribution, in the context of clustering. Density distribution has been the de facto descriptor of data distribution since the introduction of statistics. We show that density distribution has its fundamental limitation -- high-density bias, irrespective of the algorithms used to perform clustering. Existing density-based clustering algorithms have employed different algorithmic means to counter the effect of the high-density bias with some success, but the fundamental limitation of using density distribution remains an obstacle to discovering clusters of arbitrary shapes, sizes and densities. Using the mass distribution as a better foundation, we propose a new algorithm which maximizes the total mass of all clusters, called mass-maximization clustering (MMC). The algorithm can be easily changed to maximize the total density of all clusters in order to examine the fundamental limitation of using density distribution versus mass distribution. The key advantage of the MMC over the density-maximization clustering is that the maximization is conducted without a bias towards dense clusters.
title Mass Distribution versus Density Distribution in the Context of Clustering
topic Machine Learning
url https://arxiv.org/abs/2601.10759