Large-Sample Bayesian Approximations for Privatized Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Awan, Jordan, Chen, Xi, Molinari, Roberto
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910172186673152
author Awan, Jordan
Chen, Xi
Molinari, Roberto
author_facet Awan, Jordan
Chen, Xi
Molinari, Roberto
contents The increased use of differential privacy (DP) has allowed the sharing of large amounts of data while reducing the risk of disclosure of sensitive information at the individual level. However, the noise introduced by DP methods makes performing statistical inference more challenging. While various methods have been proposed to address different inferential tasks, they often require strong parametric assumptions and/or do not scale well with sample sizes (e.g. U.S. Census products). In response to these limitations, we propose an approximate Bayesian method to analyze privatized data products, which uses a two-step approach of imputing the confidential data and then sampling from the non-private posterior, and which is inspired by the method of Guha and Reiter (2025). We prove that this approximate sampler is asymptotically valid under mild assumptions. While this approach is motivated by Bayesian theory, we show through simulations that it provides conservative frequentist properties as well. We demonstrate the utility of our method by applying it in simulated settings as well as for an analysis on the drivers of homeownership via the 2022 American Community Survey.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24817
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large-Sample Bayesian Approximations for Privatized Data
Awan, Jordan
Chen, Xi
Molinari, Roberto
Methodology
Statistics Theory
Applications
The increased use of differential privacy (DP) has allowed the sharing of large amounts of data while reducing the risk of disclosure of sensitive information at the individual level. However, the noise introduced by DP methods makes performing statistical inference more challenging. While various methods have been proposed to address different inferential tasks, they often require strong parametric assumptions and/or do not scale well with sample sizes (e.g. U.S. Census products). In response to these limitations, we propose an approximate Bayesian method to analyze privatized data products, which uses a two-step approach of imputing the confidential data and then sampling from the non-private posterior, and which is inspired by the method of Guha and Reiter (2025). We prove that this approximate sampler is asymptotically valid under mild assumptions. While this approach is motivated by Bayesian theory, we show through simulations that it provides conservative frequentist properties as well. We demonstrate the utility of our method by applying it in simulated settings as well as for an analysis on the drivers of homeownership via the 2022 American Community Survey.
title Large-Sample Bayesian Approximations for Privatized Data
topic Methodology
Statistics Theory
Applications
url https://arxiv.org/abs/2604.24817