sanba: An R Package for Bayesian Clustering of Distributions via Shared Atoms Nested Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Denti, Francesco, D'Angelo, Laura
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908487828635648
author Denti, Francesco
D'Angelo, Laura
author_facet Denti, Francesco
D'Angelo, Laura
contents Nested data structures arise when observations are grouped into distinct units, such as patients within hospitals or students within schools. Accounting for this hierarchical organization is essential for valid inference, as ignoring it can lead to biased estimates and poor generalization. This article addresses the challenge of clustering both individual observations and their corresponding groups while flexibly estimating group-specific densities. Bayesian nested mixture models offer a principled and robust framework for this task. However, their practical use has often been limited by computational complexity. To overcome this barrier, we present sanba, an R package for Bayesian analysis of grouped data using nested mixture models with a shared set of atoms, a structure recently introduced in the statistical literature. The package provides multiple inference strategies, including state-of-the-art Markov Chain Monte Carlo routines and variational inference algorithms tailored for large-scale datasets. All core functions are implemented in C++ and seamlessly integrated into R, making sanba a fast and user-friendly tool for fitting nested mixture models with modern Bayesian algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle sanba: An R Package for Bayesian Clustering of Distributions via Shared Atoms Nested Models
Denti, Francesco
D'Angelo, Laura
Computation
Nested data structures arise when observations are grouped into distinct units, such as patients within hospitals or students within schools. Accounting for this hierarchical organization is essential for valid inference, as ignoring it can lead to biased estimates and poor generalization. This article addresses the challenge of clustering both individual observations and their corresponding groups while flexibly estimating group-specific densities. Bayesian nested mixture models offer a principled and robust framework for this task. However, their practical use has often been limited by computational complexity. To overcome this barrier, we present sanba, an R package for Bayesian analysis of grouped data using nested mixture models with a shared set of atoms, a structure recently introduced in the statistical literature. The package provides multiple inference strategies, including state-of-the-art Markov Chain Monte Carlo routines and variational inference algorithms tailored for large-scale datasets. All core functions are implemented in C++ and seamlessly integrated into R, making sanba a fast and user-friendly tool for fitting nested mixture models with modern Bayesian algorithms.
title sanba: An R Package for Bayesian Clustering of Distributions via Shared Atoms Nested Models
topic Computation
url https://arxiv.org/abs/2508.09758