Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910584579031040 |
|---|---|
| author | Rashid, Md Rafi Ur Liu, Jing Koike-Akino, Toshiaki Mehnaz, Shagufta Wang, Ye |
| author_facet | Rashid, Md Rafi Ur Liu, Jing Koike-Akino, Toshiaki Mehnaz, Shagufta Wang, Ye |
| contents | Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish without rigorous verification. This scenario creates a privacy threat, as pre-trained models can be intentionally crafted to compromise the privacy of fine-tuning datasets. In this study, we introduce a novel poisoning technique that uses model-unlearning as an attack tool. This approach manipulates a pre-trained language model to increase the leakage of private data during the fine-tuning process. Our method enhances both membership inference and data extraction attacks while preserving model utility. Experimental results across different models, datasets, and fine-tuning setups demonstrate that our attacks significantly surpass baseline performance. This work serves as a cautionary note for users who download pre-trained models from unverified sources, highlighting the potential risks involved. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_17354 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage Rashid, Md Rafi Ur Liu, Jing Koike-Akino, Toshiaki Mehnaz, Shagufta Wang, Ye Machine Learning Artificial Intelligence Cryptography and Security Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish without rigorous verification. This scenario creates a privacy threat, as pre-trained models can be intentionally crafted to compromise the privacy of fine-tuning datasets. In this study, we introduce a novel poisoning technique that uses model-unlearning as an attack tool. This approach manipulates a pre-trained language model to increase the leakage of private data during the fine-tuning process. Our method enhances both membership inference and data extraction attacks while preserving model utility. Experimental results across different models, datasets, and fine-tuning setups demonstrate that our attacks significantly surpass baseline performance. This work serves as a cautionary note for users who download pre-trained models from unverified sources, highlighting the potential risks involved. |
| title | Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage |
| topic | Machine Learning Artificial Intelligence Cryptography and Security |
| url | https://arxiv.org/abs/2408.17354 |