PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yidan, Cao, Yanan, Ren, Yubing, Fang, Fang, Lin, Zheng, Fang, Binxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909611908399104
author Wang, Yidan
Cao, Yanan
Ren, Yubing
Fang, Fang
Lin, Zheng
Fang, Binxing
author_facet Wang, Yidan
Cao, Yanan
Ren, Yubing
Fang, Fang
Lin, Zheng
Fang, Binxing
contents Large Language Models (LLMs) excel in various domains but pose inherent privacy risks. Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment models can easily block. Meanwhile, Jailbreak attacks bypass LLM safety mechanisms to generate harmful content, but their role in privacy scenarios remains underexplored. In this paper, we examine the effectiveness of jailbreak attacks in extracting sensitive information, bridging privacy leakage and jailbreak attacks in LLMs. Moreover, we propose PIG, a novel framework targeting Personally Identifiable Information (PII) and addressing the limitations of current jailbreak methods. Specifically, PIG identifies PII entities and their types in privacy queries, uses in-context learning to build a privacy context, and iteratively updates it with three gradient-based strategies to elicit target PII. We evaluate PIG and existing jailbreak methods using two privacy-related datasets. Experiments on four white-box and two black-box LLMs show that PIG outperforms baseline methods and achieves state-of-the-art (SoTA) results. The results underscore significant privacy risks in LLMs, emphasizing the need for stronger safeguards. Our code is availble at https://github.com/redwyd/PrivacyJailbreak.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
Wang, Yidan
Cao, Yanan
Ren, Yubing
Fang, Fang
Lin, Zheng
Fang, Binxing
Cryptography and Security
Computation and Language
Large Language Models (LLMs) excel in various domains but pose inherent privacy risks. Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment models can easily block. Meanwhile, Jailbreak attacks bypass LLM safety mechanisms to generate harmful content, but their role in privacy scenarios remains underexplored. In this paper, we examine the effectiveness of jailbreak attacks in extracting sensitive information, bridging privacy leakage and jailbreak attacks in LLMs. Moreover, we propose PIG, a novel framework targeting Personally Identifiable Information (PII) and addressing the limitations of current jailbreak methods. Specifically, PIG identifies PII entities and their types in privacy queries, uses in-context learning to build a privacy context, and iteratively updates it with three gradient-based strategies to elicit target PII. We evaluate PIG and existing jailbreak methods using two privacy-related datasets. Experiments on four white-box and two black-box LLMs show that PIG outperforms baseline methods and achieves state-of-the-art (SoTA) results. The results underscore significant privacy risks in LLMs, emphasizing the need for stronger safeguards. Our code is availble at https://github.com/redwyd/PrivacyJailbreak.
title PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2505.09921