Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rashid, Md Rafi Ur, Liu, Jing, Koike-Akino, Toshiaki, Mehnaz, Shagufta, Wang, Ye
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910584579031040
author Rashid, Md Rafi Ur
Liu, Jing
Koike-Akino, Toshiaki
Mehnaz, Shagufta
Wang, Ye
author_facet Rashid, Md Rafi Ur
Liu, Jing
Koike-Akino, Toshiaki
Mehnaz, Shagufta
Wang, Ye
contents Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish without rigorous verification. This scenario creates a privacy threat, as pre-trained models can be intentionally crafted to compromise the privacy of fine-tuning datasets. In this study, we introduce a novel poisoning technique that uses model-unlearning as an attack tool. This approach manipulates a pre-trained language model to increase the leakage of private data during the fine-tuning process. Our method enhances both membership inference and data extraction attacks while preserving model utility. Experimental results across different models, datasets, and fine-tuning setups demonstrate that our attacks significantly surpass baseline performance. This work serves as a cautionary note for users who download pre-trained models from unverified sources, highlighting the potential risks involved.
format Preprint
id arxiv_https___arxiv_org_abs_2408_17354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
Rashid, Md Rafi Ur
Liu, Jing
Koike-Akino, Toshiaki
Mehnaz, Shagufta
Wang, Ye
Machine Learning
Artificial Intelligence
Cryptography and Security
Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish without rigorous verification. This scenario creates a privacy threat, as pre-trained models can be intentionally crafted to compromise the privacy of fine-tuning datasets. In this study, we introduce a novel poisoning technique that uses model-unlearning as an attack tool. This approach manipulates a pre-trained language model to increase the leakage of private data during the fine-tuning process. Our method enhances both membership inference and data extraction attacks while preserving model utility. Experimental results across different models, datasets, and fine-tuning setups demonstrate that our attacks significantly surpass baseline performance. This work serves as a cautionary note for users who download pre-trained models from unverified sources, highlighting the potential risks involved.
title Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2408.17354