Nonparametric Inference on Unlabeled Histograms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Yun, Yang, Pengkun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909892354244608
author Ma, Yun
Yang, Pengkun
author_facet Ma, Yun
Yang, Pengkun
contents Statistical inference on histograms and frequency counts plays a central role in categorical data analysis. Moving beyond classical methods that directly analyze labeled frequencies, we introduce a framework that models the multiset of unlabeled histograms via a mixture distribution to better capture unseen domain elements in large-alphabet regime. We study the nonparametric maximum likelihood estimator (NPMLE) under this framework, and establish its optimal convergence rate under the Poisson setting. The NPMLE also immediately yields flexible and efficient plug-in estimators for functional estimation problems, where a localized variant further achieves the optimal sample complexity for a wide range of symmetric functionals. Extensive experiments on synthetic, real-world datasets, and large language models highlight the practical benefits of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Nonparametric Inference on Unlabeled Histograms
Ma, Yun
Yang, Pengkun
Statistics Theory
Methodology
Statistical inference on histograms and frequency counts plays a central role in categorical data analysis. Moving beyond classical methods that directly analyze labeled frequencies, we introduce a framework that models the multiset of unlabeled histograms via a mixture distribution to better capture unseen domain elements in large-alphabet regime. We study the nonparametric maximum likelihood estimator (NPMLE) under this framework, and establish its optimal convergence rate under the Poisson setting. The NPMLE also immediately yields flexible and efficient plug-in estimators for functional estimation problems, where a localized variant further achieves the optimal sample complexity for a wide range of symmetric functionals. Extensive experiments on synthetic, real-world datasets, and large language models highlight the practical benefits of the proposed method.
title Nonparametric Inference on Unlabeled Histograms
topic Statistics Theory
Methodology
url https://arxiv.org/abs/2511.05077