Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Yuxin, Marchyok, Leo, Hong, Sanghyun, Geiping, Jonas, Goldstein, Tom, Carlini, Nicholas
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911821616644096
author Wen, Yuxin
Marchyok, Leo
Hong, Sanghyun
Geiping, Jonas
Goldstein, Tom
Carlini, Nicholas
author_facet Wen, Yuxin
Marchyok, Leo
Hong, Sanghyun
Geiping, Jonas
Goldstein, Tom
Carlini, Nicholas
contents It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model checkpoints on the web poses considerable risks, including the vulnerability to backdoor attacks. In this paper, we unveil a new vulnerability: the privacy backdoor attack. This black-box privacy attack aims to amplify the privacy leakage that arises when fine-tuning a model: when a victim fine-tunes a backdoored model, their training data will be leaked at a significantly higher rate than if they had fine-tuned a typical model. We conduct extensive experiments on various datasets and models, including both vision-language models (CLIP) and large language models, demonstrating the broad applicability and effectiveness of such an attack. Additionally, we carry out multiple ablation studies with different fine-tuning methods and inference strategies to thoroughly analyze this new threat. Our findings highlight a critical privacy concern within the machine learning community and call for a reevaluation of safety protocols in the use of open-source pre-trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01231
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
Wen, Yuxin
Marchyok, Leo
Hong, Sanghyun
Geiping, Jonas
Goldstein, Tom
Carlini, Nicholas
Cryptography and Security
Machine Learning
It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model checkpoints on the web poses considerable risks, including the vulnerability to backdoor attacks. In this paper, we unveil a new vulnerability: the privacy backdoor attack. This black-box privacy attack aims to amplify the privacy leakage that arises when fine-tuning a model: when a victim fine-tunes a backdoored model, their training data will be leaked at a significantly higher rate than if they had fine-tuned a typical model. We conduct extensive experiments on various datasets and models, including both vision-language models (CLIP) and large language models, demonstrating the broad applicability and effectiveness of such an attack. Additionally, we carry out multiple ablation studies with different fine-tuning methods and inference strategies to thoroughly analyze this new threat. Our findings highlight a critical privacy concern within the machine learning community and call for a reevaluation of safety protocols in the use of open-source pre-trained models.
title Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2404.01231