Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chua, Lynn, Ghazi, Badih, Huang, Yangsibo, Kamath, Pritish, Kumar, Ravi, Liu, Daogao, Manurangsi, Pasin, Sinha, Amer, Zhang, Chiyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916359735083008
author Chua, Lynn
Ghazi, Badih
Huang, Yangsibo
Kamath, Pritish
Kumar, Ravi
Liu, Daogao
Manurangsi, Pasin
Sinha, Amer
Zhang, Chiyuan
author_facet Chua, Lynn
Ghazi, Badih
Huang, Yangsibo
Kamath, Pritish
Kumar, Ravi
Liu, Daogao
Manurangsi, Pasin
Sinha, Amer
Zhang, Chiyuan
contents Large language models (LLMs) have emerged as powerful tools for tackling complex tasks across diverse domains, but they also raise privacy concerns when fine-tuned on sensitive data due to potential memorization. While differential privacy (DP) offers a promising solution by ensuring models are 'almost indistinguishable' with or without any particular privacy unit, current evaluations on LLMs mostly treat each example (text record) as the privacy unit. This leads to uneven user privacy guarantees when contributions per user vary. We therefore study user-level DP motivated by applications where it necessary to ensure uniform privacy protection across users. We present a systematic evaluation of user-level DP for LLM fine-tuning on natural language generation tasks. Focusing on two mechanisms for achieving user-level DP guarantees, Group Privacy and User-wise DP-SGD, we investigate design choices like data selection strategies and parameter tuning for the best privacy-utility tradeoff.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14322
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
Chua, Lynn
Ghazi, Badih
Huang, Yangsibo
Kamath, Pritish
Kumar, Ravi
Liu, Daogao
Manurangsi, Pasin
Sinha, Amer
Zhang, Chiyuan
Computation and Language
Cryptography and Security
Machine Learning
Large language models (LLMs) have emerged as powerful tools for tackling complex tasks across diverse domains, but they also raise privacy concerns when fine-tuned on sensitive data due to potential memorization. While differential privacy (DP) offers a promising solution by ensuring models are 'almost indistinguishable' with or without any particular privacy unit, current evaluations on LLMs mostly treat each example (text record) as the privacy unit. This leads to uneven user privacy guarantees when contributions per user vary. We therefore study user-level DP motivated by applications where it necessary to ensure uniform privacy protection across users. We present a systematic evaluation of user-level DP for LLM fine-tuning on natural language generation tasks. Focusing on two mechanisms for achieving user-level DP guarantees, Group Privacy and User-wise DP-SGD, we investigate design choices like data selection strategies and parameter tuning for the best privacy-utility tradeoff.
title Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2406.14322