Private Estimation when Data and Privacy Demands are Correlated

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhuri, Syomantak, Courtade, Thomas A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910914656075776
author Chaudhuri, Syomantak
Courtade, Thomas A.
author_facet Chaudhuri, Syomantak
Courtade, Thomas A.
contents Differential Privacy (DP) is the current gold-standard for ensuring privacy for statistical queries. Estimation problems under DP constraints appearing in the literature have largely focused on providing equal privacy to all users. We consider the problems of empirical mean estimation for univariate data and frequency estimation for categorical data, both subject to heterogeneous privacy constraints. Each user, contributing a sample to the dataset, is allowed to have a different privacy demand. The dataset itself is assumed to be worst-case and we study both problems under two different formulations -- first, where privacy demands and data may be correlated, and second, where correlations are weakened by random permutation of the dataset. We establish theoretical performance guarantees for our proposed algorithms, under both PAC error and mean-squared error. These performance guarantees translate to minimax optimality in several instances, and experiments confirm superior performance of our algorithms over other baseline techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Private Estimation when Data and Privacy Demands are Correlated
Chaudhuri, Syomantak
Courtade, Thomas A.
Machine Learning
Cryptography and Security
Differential Privacy (DP) is the current gold-standard for ensuring privacy for statistical queries. Estimation problems under DP constraints appearing in the literature have largely focused on providing equal privacy to all users. We consider the problems of empirical mean estimation for univariate data and frequency estimation for categorical data, both subject to heterogeneous privacy constraints. Each user, contributing a sample to the dataset, is allowed to have a different privacy demand. The dataset itself is assumed to be worst-case and we study both problems under two different formulations -- first, where privacy demands and data may be correlated, and second, where correlations are weakened by random permutation of the dataset. We establish theoretical performance guarantees for our proposed algorithms, under both PAC error and mean-squared error. These performance guarantees translate to minimax optimality in several instances, and experiments confirm superior performance of our algorithms over other baseline techniques.
title Private Estimation when Data and Privacy Demands are Correlated
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2407.11274