Generalized Data Thinning Using Sufficient Statistics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dharamshi, Ameer, Neufeld, Anna, Motwani, Keshav, Gao, Lucy L., Witten, Daniela, Bien, Jacob
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914211235364864
author Dharamshi, Ameer
Neufeld, Anna
Motwani, Keshav
Gao, Lucy L.
Witten, Daniela
Bien, Jacob
author_facet Dharamshi, Ameer
Neufeld, Anna
Motwani, Keshav
Gao, Lucy L.
Witten, Daniela
Bien, Jacob
contents Our goal is to develop a general strategy to decompose a random variable $X$ into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some well-known natural exponential families, $X$ can be "thinned" into independent random variables $X^{(1)}, \ldots, X^{(K)}$, such that $X = \sum_{k=1}^K X^{(k)}$. These independent random variables can then be used for various model validation and inference tasks, including in contexts where traditional sample splitting fails. In this paper, we generalize their procedure by relaxing this summation requirement and simply asking that some known function of the independent random variables exactly reconstruct $X$. This generalization of the procedure serves two purposes. First, it greatly expands the families of distributions for which thinning can be performed. Second, it unifies sample splitting and data thinning, which on the surface seem to be very different, as applications of the same principle. This shared principle is sufficiency. We use this insight to perform generalized thinning operations for a diverse set of families.
format Preprint
id arxiv_https___arxiv_org_abs_2303_12931
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generalized Data Thinning Using Sufficient Statistics
Dharamshi, Ameer
Neufeld, Anna
Motwani, Keshav
Gao, Lucy L.
Witten, Daniela
Bien, Jacob
Methodology
Statistics Theory
Machine Learning
Our goal is to develop a general strategy to decompose a random variable $X$ into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some well-known natural exponential families, $X$ can be "thinned" into independent random variables $X^{(1)}, \ldots, X^{(K)}$, such that $X = \sum_{k=1}^K X^{(k)}$. These independent random variables can then be used for various model validation and inference tasks, including in contexts where traditional sample splitting fails. In this paper, we generalize their procedure by relaxing this summation requirement and simply asking that some known function of the independent random variables exactly reconstruct $X$. This generalization of the procedure serves two purposes. First, it greatly expands the families of distributions for which thinning can be performed. Second, it unifies sample splitting and data thinning, which on the surface seem to be very different, as applications of the same principle. This shared principle is sufficiency. We use this insight to perform generalized thinning operations for a diverse set of families.
title Generalized Data Thinning Using Sufficient Statistics
topic Methodology
Statistics Theory
Machine Learning
url https://arxiv.org/abs/2303.12931