Clustering risk in Non-parametric Hidden Markov and I.I.D. Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gassiat, Elisabeth, Kaddouri, Ibrahim, Naulet, Zacharie
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915306318856192
author Gassiat, Elisabeth
Kaddouri, Ibrahim
Naulet, Zacharie
author_facet Gassiat, Elisabeth
Kaddouri, Ibrahim
Naulet, Zacharie
contents We conduct an in-depth analysis of the Bayes risk of clustering in the context of Hidden Markov and i.i.d. models. In both settings, we identify the situations where this risk is comparable to the Bayes risk of classification and those where its minimizer, the Bayes clusterer, can be derived from the Bayes classifier. While we demonstrate that clustering based on the Bayes classifier does not always match the optimal Bayes clusterer, we show that this difference is primarily theoretical and that the Bayes classifier remains nearly optimal for clustering. A key quantity emerges, capturing the fundamental difficulty of both classification and clustering tasks. Furthermore, by leveraging the identifiability of HMMs, we establish bounds on the clustering excess risk of a plug-in Bayes classifier in the general nonparametric setting, offering theoretical justification for its widespread use in practice. Simulations further illustrate our findings.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12238
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Clustering risk in Non-parametric Hidden Markov and I.I.D. Models
Gassiat, Elisabeth
Kaddouri, Ibrahim
Naulet, Zacharie
Statistics Theory
Machine Learning
62M05, 62G99
We conduct an in-depth analysis of the Bayes risk of clustering in the context of Hidden Markov and i.i.d. models. In both settings, we identify the situations where this risk is comparable to the Bayes risk of classification and those where its minimizer, the Bayes clusterer, can be derived from the Bayes classifier. While we demonstrate that clustering based on the Bayes classifier does not always match the optimal Bayes clusterer, we show that this difference is primarily theoretical and that the Bayes classifier remains nearly optimal for clustering. A key quantity emerges, capturing the fundamental difficulty of both classification and clustering tasks. Furthermore, by leveraging the identifiability of HMMs, we establish bounds on the clustering excess risk of a plug-in Bayes classifier in the general nonparametric setting, offering theoretical justification for its widespread use in practice. Simulations further illustrate our findings.
title Clustering risk in Non-parametric Hidden Markov and I.I.D. Models
topic Statistics Theory
Machine Learning
62M05, 62G99
url https://arxiv.org/abs/2309.12238