Measures of Overlapping Multivariate Gaussian Clusters in Unsupervised Online Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ožbot, Miha, Škrjanc, Igor
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912546693316608
author Ožbot, Miha
Škrjanc, Igor
author_facet Ožbot, Miha
Škrjanc, Igor
contents In this paper, we propose a new measure for detecting overlap in multivariate Gaussian clusters. The aim of online learning from data streams is to create clustering, classification, or regression models that can adapt over time based on the conceptual drift of streaming data. In the case of clustering, this can result in a large number of clusters that may overlap and should be merged. Commonly used distribution dissimilarity measures are not adequate for determining overlapping clusters in the context of online learning from streaming data due to their inability to account for all shapes of clusters and their high computational demands. Our proposed dissimilarity measure is specifically designed to detect overlap rather than dissimilarity and can be computed faster compared to existing measures. Our method is several times faster than compared methods and is capable of detecting overlapping clusters while avoiding the merging of orthogonal clusters.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15444
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measures of Overlapping Multivariate Gaussian Clusters in Unsupervised Online Learning
Ožbot, Miha
Škrjanc, Igor
Machine Learning
68T05, 62H30, 62H20
I.2.6; I.5.3
In this paper, we propose a new measure for detecting overlap in multivariate Gaussian clusters. The aim of online learning from data streams is to create clustering, classification, or regression models that can adapt over time based on the conceptual drift of streaming data. In the case of clustering, this can result in a large number of clusters that may overlap and should be merged. Commonly used distribution dissimilarity measures are not adequate for determining overlapping clusters in the context of online learning from streaming data due to their inability to account for all shapes of clusters and their high computational demands. Our proposed dissimilarity measure is specifically designed to detect overlap rather than dissimilarity and can be computed faster compared to existing measures. Our method is several times faster than compared methods and is capable of detecting overlapping clusters while avoiding the merging of orthogonal clusters.
title Measures of Overlapping Multivariate Gaussian Clusters in Unsupervised Online Learning
topic Machine Learning
68T05, 62H30, 62H20
I.2.6; I.5.3
url https://arxiv.org/abs/2508.15444