Doubly Stochastic Mean-Shift Clustering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Trigano, Tom, Sepulcre, Yann, Lapidot, Itshak
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914334222843904
author Trigano, Tom
Sepulcre, Yann
Lapidot, Itshak
author_facet Trigano, Tom
Sepulcre, Yann
Lapidot, Itshak
contents Standard Mean-Shift algorithms are notoriously sensitive to the bandwidth hyperparameter, particularly in data-scarce regimes where fixed-scale density estimation leads to fragmentation and spurious modes. In this paper, we propose Doubly Stochastic Mean-Shift (DSMS), a novel extension that introduces randomness not only in the trajectory updates but also in the kernel bandwidth itself. By drawing both the data samples and the radius from a continuous uniform distribution at each iteration, DSMS effectively performs a better exploration of the density landscape. We show that this randomized bandwidth policy acts as an implicit regularization mechanism, and provide convergence theoretical results. Comparative experiments on synthetic Gaussian mixtures reveal that DSMS significantly outperforms standard and stochastic Mean-Shift baselines, exhibiting remarkable stability and preventing over-segmentation in sparse clustering scenarios without other performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15393
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Doubly Stochastic Mean-Shift Clustering
Trigano, Tom
Sepulcre, Yann
Lapidot, Itshak
Machine Learning
Computer Vision and Pattern Recognition
Standard Mean-Shift algorithms are notoriously sensitive to the bandwidth hyperparameter, particularly in data-scarce regimes where fixed-scale density estimation leads to fragmentation and spurious modes. In this paper, we propose Doubly Stochastic Mean-Shift (DSMS), a novel extension that introduces randomness not only in the trajectory updates but also in the kernel bandwidth itself. By drawing both the data samples and the radius from a continuous uniform distribution at each iteration, DSMS effectively performs a better exploration of the density landscape. We show that this randomized bandwidth policy acts as an implicit regularization mechanism, and provide convergence theoretical results. Comparative experiments on synthetic Gaussian mixtures reveal that DSMS significantly outperforms standard and stochastic Mean-Shift baselines, exhibiting remarkable stability and preventing over-segmentation in sparse clustering scenarios without other performance degradation.
title Doubly Stochastic Mean-Shift Clustering
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.15393