Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Xinyu, Shin, Richard, Inan, Huseyin A., Manoel, Andre, Mireshghallah, Fatemehsadat, Lin, Zinan, Gopi, Sivakanth, Kulkarni, Janardhan, Sim, Robert
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913212129083392
author Tang, Xinyu
Shin, Richard
Inan, Huseyin A.
Manoel, Andre
Mireshghallah, Fatemehsadat
Lin, Zinan
Gopi, Sivakanth
Kulkarni, Janardhan
Sim, Robert
author_facet Tang, Xinyu
Shin, Richard
Inan, Huseyin A.
Manoel, Andre
Mireshghallah, Fatemehsadat
Lin, Zinan
Gopi, Sivakanth
Kulkarni, Janardhan
Sim, Robert
contents We study the problem of in-context learning (ICL) with large language models (LLMs) on private datasets. This scenario poses privacy risks, as LLMs may leak or regurgitate the private examples demonstrated in the prompt. We propose a novel algorithm that generates synthetic few-shot demonstrations from the private dataset with formal differential privacy (DP) guarantees, and show empirically that it can achieve effective ICL. We conduct extensive experiments on standard benchmarks and compare our algorithm with non-private ICL and zero-shot solutions. Our results demonstrate that our algorithm can achieve competitive performance with strong privacy levels. These results open up new possibilities for ICL with privacy protection for a broad range of applications.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11765
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
Tang, Xinyu
Shin, Richard
Inan, Huseyin A.
Manoel, Andre
Mireshghallah, Fatemehsadat
Lin, Zinan
Gopi, Sivakanth
Kulkarni, Janardhan
Sim, Robert
Machine Learning
Cryptography and Security
We study the problem of in-context learning (ICL) with large language models (LLMs) on private datasets. This scenario poses privacy risks, as LLMs may leak or regurgitate the private examples demonstrated in the prompt. We propose a novel algorithm that generates synthetic few-shot demonstrations from the private dataset with formal differential privacy (DP) guarantees, and show empirically that it can achieve effective ICL. We conduct extensive experiments on standard benchmarks and compare our algorithm with non-private ICL and zero-shot solutions. Our results demonstrate that our algorithm can achieve competitive performance with strong privacy levels. These results open up new possibilities for ICL with privacy protection for a broad range of applications.
title Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2309.11765