Differentially Private Stream Processing at Scale

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Bing, Doroshenko, Vadym, Kairouz, Peter, Steinke, Thomas, Thakurta, Abhradeep, Ma, Ziyin, Cohen, Eidan, Apte, Himani, Spacek, Jodi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914868114751488
author Zhang, Bing
Doroshenko, Vadym
Kairouz, Peter
Steinke, Thomas
Thakurta, Abhradeep
Ma, Ziyin
Cohen, Eidan
Apte, Himani
Spacek, Jodi
author_facet Zhang, Bing
Doroshenko, Vadym
Kairouz, Peter
Steinke, Thomas
Thakurta, Abhradeep
Ma, Ziyin
Cohen, Eidan
Apte, Himani
Spacek, Jodi
contents We design, to the best of our knowledge, the first differentially private (DP) stream aggregation processing system at scale. Our system -- Differential Privacy SQL Pipelines (DP-SQLP) -- is built using a streaming framework similar to Spark streaming, and is built on top of the Spanner database and the F1 query engine from Google. Towards designing DP-SQLP we make both algorithmic and systemic advances, namely, we (i) design a novel (user-level) DP key selection algorithm that can operate on an unbounded set of possible keys, and can scale to one billion keys that users have contributed, (ii) design a preemptive execution scheme for DP key selection that avoids enumerating all the keys at each triggering time, and (iii) use algorithmic techniques from DP continual observation to release a continual DP histogram of user contributions to different keys over the stream length. We empirically demonstrate the efficacy by obtaining at least $16\times$ reduction in error over meaningful baselines we consider. We implemented a streaming differentially private user impressions for Google Shopping with DP-SQLP. The streaming DP algorithms are further applied to Google Trends.
format Preprint
id arxiv_https___arxiv_org_abs_2303_18086
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Differentially Private Stream Processing at Scale
Zhang, Bing
Doroshenko, Vadym
Kairouz, Peter
Steinke, Thomas
Thakurta, Abhradeep
Ma, Ziyin
Cohen, Eidan
Apte, Himani
Spacek, Jodi
Cryptography and Security
Databases
We design, to the best of our knowledge, the first differentially private (DP) stream aggregation processing system at scale. Our system -- Differential Privacy SQL Pipelines (DP-SQLP) -- is built using a streaming framework similar to Spark streaming, and is built on top of the Spanner database and the F1 query engine from Google. Towards designing DP-SQLP we make both algorithmic and systemic advances, namely, we (i) design a novel (user-level) DP key selection algorithm that can operate on an unbounded set of possible keys, and can scale to one billion keys that users have contributed, (ii) design a preemptive execution scheme for DP key selection that avoids enumerating all the keys at each triggering time, and (iii) use algorithmic techniques from DP continual observation to release a continual DP histogram of user contributions to different keys over the stream length. We empirically demonstrate the efficacy by obtaining at least $16\times$ reduction in error over meaningful baselines we consider. We implemented a streaming differentially private user impressions for Google Shopping with DP-SQLP. The streaming DP algorithms are further applied to Google Trends.
title Differentially Private Stream Processing at Scale
topic Cryptography and Security
Databases
url https://arxiv.org/abs/2303.18086