Privacy Side Channels in Machine Learning Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Debenedetti, Edoardo, Severi, Giorgio, Carlini, Nicholas, Choquette-Choo, Christopher A., Jagielski, Matthew, Nasr, Milad, Wallace, Eric, Tramèr, Florian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909260968886272
author Debenedetti, Edoardo
Severi, Giorgio
Carlini, Nicholas
Choquette-Choo, Christopher A.
Jagielski, Matthew
Nasr, Milad
Wallace, Eric
Tramèr, Florian
author_facet Debenedetti, Edoardo
Severi, Giorgio
Carlini, Nicholas
Choquette-Choo, Christopher A.
Jagielski, Matthew
Nasr, Milad
Wallace, Eric
Tramèr, Florian
contents Most current approaches for protecting privacy in machine learning (ML) assume that models exist in a vacuum. Yet, in reality, these models are part of larger systems that include components for training data filtering, output monitoring, and more. In this work, we introduce privacy side channels: attacks that exploit these system-level components to extract private information at far higher rates than is otherwise possible for standalone models. We propose four categories of side channels that span the entire ML lifecycle (training data filtering, input preprocessing, output post-processing, and query filtering) and allow for enhanced membership inference, data extraction, and even novel threats such as extraction of users' test queries. For example, we show that deduplicating training data before applying differentially-private training creates a side-channel that completely invalidates any provable privacy guarantees. We further show that systems which block language models from regenerating training data can be exploited to exfiltrate private keys contained in the training set--even if the model did not memorize these keys. Taken together, our results demonstrate the need for a holistic, end-to-end privacy analysis of machine learning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05610
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Privacy Side Channels in Machine Learning Systems
Debenedetti, Edoardo
Severi, Giorgio
Carlini, Nicholas
Choquette-Choo, Christopher A.
Jagielski, Matthew
Nasr, Milad
Wallace, Eric
Tramèr, Florian
Cryptography and Security
Machine Learning
Most current approaches for protecting privacy in machine learning (ML) assume that models exist in a vacuum. Yet, in reality, these models are part of larger systems that include components for training data filtering, output monitoring, and more. In this work, we introduce privacy side channels: attacks that exploit these system-level components to extract private information at far higher rates than is otherwise possible for standalone models. We propose four categories of side channels that span the entire ML lifecycle (training data filtering, input preprocessing, output post-processing, and query filtering) and allow for enhanced membership inference, data extraction, and even novel threats such as extraction of users' test queries. For example, we show that deduplicating training data before applying differentially-private training creates a side-channel that completely invalidates any provable privacy guarantees. We further show that systems which block language models from regenerating training data can be exploited to exfiltrate private keys contained in the training set--even if the model did not memorize these keys. Taken together, our results demonstrate the need for a holistic, end-to-end privacy analysis of machine learning systems.
title Privacy Side Channels in Machine Learning Systems
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2309.05610