Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911821616644096 |
|---|---|
| author | Wen, Yuxin Marchyok, Leo Hong, Sanghyun Geiping, Jonas Goldstein, Tom Carlini, Nicholas |
| author_facet | Wen, Yuxin Marchyok, Leo Hong, Sanghyun Geiping, Jonas Goldstein, Tom Carlini, Nicholas |
| contents | It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model checkpoints on the web poses considerable risks, including the vulnerability to backdoor attacks. In this paper, we unveil a new vulnerability: the privacy backdoor attack. This black-box privacy attack aims to amplify the privacy leakage that arises when fine-tuning a model: when a victim fine-tunes a backdoored model, their training data will be leaked at a significantly higher rate than if they had fine-tuned a typical model. We conduct extensive experiments on various datasets and models, including both vision-language models (CLIP) and large language models, demonstrating the broad applicability and effectiveness of such an attack. Additionally, we carry out multiple ablation studies with different fine-tuning methods and inference strategies to thoroughly analyze this new threat. Our findings highlight a critical privacy concern within the machine learning community and call for a reevaluation of safety protocols in the use of open-source pre-trained models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_01231 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models Wen, Yuxin Marchyok, Leo Hong, Sanghyun Geiping, Jonas Goldstein, Tom Carlini, Nicholas Cryptography and Security Machine Learning It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model checkpoints on the web poses considerable risks, including the vulnerability to backdoor attacks. In this paper, we unveil a new vulnerability: the privacy backdoor attack. This black-box privacy attack aims to amplify the privacy leakage that arises when fine-tuning a model: when a victim fine-tunes a backdoored model, their training data will be leaked at a significantly higher rate than if they had fine-tuned a typical model. We conduct extensive experiments on various datasets and models, including both vision-language models (CLIP) and large language models, demonstrating the broad applicability and effectiveness of such an attack. Additionally, we carry out multiple ablation studies with different fine-tuning methods and inference strategies to thoroughly analyze this new threat. Our findings highlight a critical privacy concern within the machine learning community and call for a reevaluation of safety protocols in the use of open-source pre-trained models. |
| title | Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models |
| topic | Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2404.01231 |