On Misspecified Error Distributions in Bayesian Functional Clustering: Consequences and Remedies

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Iwashige, Fumiya, Wakayama, Tomoya, Sugasawa, Shonosuke, Hashimoto, Shintaro
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908602712719360
author Iwashige, Fumiya
Wakayama, Tomoya
Sugasawa, Shonosuke
Hashimoto, Shintaro
author_facet Iwashige, Fumiya
Wakayama, Tomoya
Sugasawa, Shonosuke
Hashimoto, Shintaro
contents Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a fundamental cause of this phenomenon lies in misspecification of the error structure: errors are conventionally assumed to be independent across observed points in Bayesian functional models. Through high-dimensional clustering theory, we demonstrate that ignoring the underlying correlation leads to excess clusters regardless of the flexibility of prior distributions. Guided by this theory, we propose incorporating the underlying correlation structures via Gaussian processes and also present its scalable approximation with principled hyperparameter selection. Numerical experiments illustrate that even simple clustering based on Dirichlet processes performs well once error dependence is properly modeled.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Misspecified Error Distributions in Bayesian Functional Clustering: Consequences and Remedies
Iwashige, Fumiya
Wakayama, Tomoya
Sugasawa, Shonosuke
Hashimoto, Shintaro
Methodology
Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a fundamental cause of this phenomenon lies in misspecification of the error structure: errors are conventionally assumed to be independent across observed points in Bayesian functional models. Through high-dimensional clustering theory, we demonstrate that ignoring the underlying correlation leads to excess clusters regardless of the flexibility of prior distributions. Guided by this theory, we propose incorporating the underlying correlation structures via Gaussian processes and also present its scalable approximation with principled hyperparameter selection. Numerical experiments illustrate that even simple clustering based on Dirichlet processes performs well once error dependence is properly modeled.
title On Misspecified Error Distributions in Bayesian Functional Clustering: Consequences and Remedies
topic Methodology
url https://arxiv.org/abs/2510.17215