Parameter-Free Clustering via Self-Supervised Consensus Maximization (Extended Version)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lijun, Liu, Suyuan, Wang, Siwei, Yu, Shengju, Zhu, Xueling, Li, Miaomiao, Liu, Xinwang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911539731103744
author Zhang, Lijun
Liu, Suyuan
Wang, Siwei
Yu, Shengju
Zhu, Xueling
Li, Miaomiao
Liu, Xinwang
author_facet Zhang, Lijun
Liu, Suyuan
Wang, Siwei
Yu, Shengju
Zhu, Xueling
Li, Miaomiao
Liu, Xinwang
contents Clustering is a fundamental task in unsupervised learning, but most existing methods heavily rely on hyperparameters such as the number of clusters or other sensitive settings, limiting their applicability in real-world scenarios. To address this long-standing challenge, we propose a novel and fully parameter-free clustering framework via Self-supervised Consensus Maximization, named SCMax. Our framework performs hierarchical agglomerative clustering and cluster evaluation in a single, integrated process. At each step of agglomeration, it creates a new, structure-aware data representation through a self-supervised learning task guided by the current clustering structure. We then introduce a nearest neighbor consensus score, which measures the agreement between the nearest neighbor-based merge decisions suggested by the original representation and the self-supervised one. The moment at which consensus maximization occurs can serve as a criterion for determining the optimal number of clusters. Extensive experiments on multiple datasets demonstrate that the proposed framework outperforms existing clustering approaches designed for scenarios with an unknown number of clusters.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parameter-Free Clustering via Self-Supervised Consensus Maximization (Extended Version)
Zhang, Lijun
Liu, Suyuan
Wang, Siwei
Yu, Shengju
Zhu, Xueling
Li, Miaomiao
Liu, Xinwang
Machine Learning
Clustering is a fundamental task in unsupervised learning, but most existing methods heavily rely on hyperparameters such as the number of clusters or other sensitive settings, limiting their applicability in real-world scenarios. To address this long-standing challenge, we propose a novel and fully parameter-free clustering framework via Self-supervised Consensus Maximization, named SCMax. Our framework performs hierarchical agglomerative clustering and cluster evaluation in a single, integrated process. At each step of agglomeration, it creates a new, structure-aware data representation through a self-supervised learning task guided by the current clustering structure. We then introduce a nearest neighbor consensus score, which measures the agreement between the nearest neighbor-based merge decisions suggested by the original representation and the self-supervised one. The moment at which consensus maximization occurs can serve as a criterion for determining the optimal number of clusters. Extensive experiments on multiple datasets demonstrate that the proposed framework outperforms existing clustering approaches designed for scenarios with an unknown number of clusters.
title Parameter-Free Clustering via Self-Supervised Consensus Maximization (Extended Version)
topic Machine Learning
url https://arxiv.org/abs/2511.09211