Saved in:
Bibliographic Details
Main Authors: Breutigam, Dennis, Reischuk, Rüdiger
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.11429
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916691824345088
author Breutigam, Dennis
Reischuk, Rüdiger
author_facet Breutigam, Dennis
Reischuk, Rüdiger
contents Differential privacy (DP) considers a scenario, where an adversary has almost complete information about the entries of a database This worst-case assumption is likely to overestimate the privacy thread for an individual in real life. Statistical privacy (SP) denotes a setting where only the distribution of the database entries is known to an adversary, but not their exact values. In this case one has to analyze the interaction between noiseless privacy based on the entropy of distributions and privacy mechanisms that distort the answers of queries, which can be quite complex. A privacy mechanism often used is to take samples of the data for answering a query. This paper proves precise bounds how much different methods of sampling increase privacy in the statistical setting with respect to database size and sampling rate. They allow us to deduce when and how much sampling provides an improvement and how far this depends on the privacy parameter ε. To perform these investigations we develop a framework to model sampling techniques. For the DP setting tradeoff functions have been proposed as a finer measure for privacy compared to (ε,δ)-pairs. We apply these tools to statistical privacy with subsampling to get a comparable characterization
format Preprint
id arxiv_https___arxiv_org_abs_2504_11429
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Statistical Privacy by Subsampling
Breutigam, Dennis
Reischuk, Rüdiger
Cryptography and Security
Differential privacy (DP) considers a scenario, where an adversary has almost complete information about the entries of a database This worst-case assumption is likely to overestimate the privacy thread for an individual in real life. Statistical privacy (SP) denotes a setting where only the distribution of the database entries is known to an adversary, but not their exact values. In this case one has to analyze the interaction between noiseless privacy based on the entropy of distributions and privacy mechanisms that distort the answers of queries, which can be quite complex. A privacy mechanism often used is to take samples of the data for answering a query. This paper proves precise bounds how much different methods of sampling increase privacy in the statistical setting with respect to database size and sampling rate. They allow us to deduce when and how much sampling provides an improvement and how far this depends on the privacy parameter ε. To perform these investigations we develop a framework to model sampling techniques. For the DP setting tradeoff functions have been proposed as a finer measure for privacy compared to (ε,δ)-pairs. We apply these tools to statistical privacy with subsampling to get a comparable characterization
title Improving Statistical Privacy by Subsampling
topic Cryptography and Security
url https://arxiv.org/abs/2504.11429