Post-clustering Inference under Dependence

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: González-Delgado, Javier, Deronzier, Mathis, Cortés, Juan, Neuvial, Pierre
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912536362745856
author González-Delgado, Javier
Deronzier, Mathis
Cortés, Juan
Neuvial, Pierre
author_facet González-Delgado, Javier
Deronzier, Mathis
Cortés, Juan
Neuvial, Pierre
contents Recent work by Gao et al. (JASA 2022) has laid the foundations for post-clustering inference, establishing a theoretical framework allowing to test for differences between means of estimated clusters. Additionally, they studied the estimation of unknown parameters while controlling the selective type I error. However, their theory was developed for independent observations identically distributed as $p$-dimensional Gaussian variables, where the parameter estimation could only be performed for spherical covariance matrices. Here, we aim at extending this framework to a more convenient scenario for practical applications, where arbitrary dependence structures between observations and features are allowed. We establish sufficient conditions for extending the setting presented by Gao et al. to the general dependence framework. Moreover, we assess theoretical conditions allowing the compatible estimation of a covariance matrix. The theory is developed for hierarchical agglomerative clustering algorithms with several types of linkages, and for the $k$-means algorithm. We illustrate our method with synthetic data and real data of protein structures.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11822
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Post-clustering Inference under Dependence
González-Delgado, Javier
Deronzier, Mathis
Cortés, Juan
Neuvial, Pierre
Methodology
Statistics Theory
Applications
Recent work by Gao et al. (JASA 2022) has laid the foundations for post-clustering inference, establishing a theoretical framework allowing to test for differences between means of estimated clusters. Additionally, they studied the estimation of unknown parameters while controlling the selective type I error. However, their theory was developed for independent observations identically distributed as $p$-dimensional Gaussian variables, where the parameter estimation could only be performed for spherical covariance matrices. Here, we aim at extending this framework to a more convenient scenario for practical applications, where arbitrary dependence structures between observations and features are allowed. We establish sufficient conditions for extending the setting presented by Gao et al. to the general dependence framework. Moreover, we assess theoretical conditions allowing the compatible estimation of a covariance matrix. The theory is developed for hierarchical agglomerative clustering algorithms with several types of linkages, and for the $k$-means algorithm. We illustrate our method with synthetic data and real data of protein structures.
title Post-clustering Inference under Dependence
topic Methodology
Statistics Theory
Applications
url https://arxiv.org/abs/2310.11822