CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhuochen, Fok, Kar Wai, Thing, Vrizlynn L. L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909997368082432
author Yang, Zhuochen
Fok, Kar Wai
Thing, Vrizlynn L. L.
author_facet Yang, Zhuochen
Fok, Kar Wai
Thing, Vrizlynn L. L.
contents Large language models have gained widespread attention recently, but their potential security vulnerabilities, especially privacy leakage, are also becoming apparent. To test and evaluate for data extraction risks in LLM, we proposed CoSPED, short for Consistent Soft Prompt targeted data Extraction and Defense. We introduce several innovative components, including Dynamic Loss, Additive Loss, Common Loss, and Self Consistency Decoding Strategy, and tested to enhance the consistency of the soft prompt tuning process. Through extensive experimentation with various combinations, we achieved an extraction rate of 65.2% at a 50-token prefix comparison. Our comparisons of CoSPED with other reference works confirm our superior extraction rates. We evaluate CoSPED on more scenarios, achieving Pythia model extraction rate of 51.7% and introducing cross-model comparison. Finally, we explore defense through Rank-One Model Editing and achieve a reduction in the extraction rate to 1.6%, which proves that our analysis of extraction mechanisms can directly inform effective mitigation strategies against soft prompt-based attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
Yang, Zhuochen
Fok, Kar Wai
Thing, Vrizlynn L. L.
Cryptography and Security
Large language models have gained widespread attention recently, but their potential security vulnerabilities, especially privacy leakage, are also becoming apparent. To test and evaluate for data extraction risks in LLM, we proposed CoSPED, short for Consistent Soft Prompt targeted data Extraction and Defense. We introduce several innovative components, including Dynamic Loss, Additive Loss, Common Loss, and Self Consistency Decoding Strategy, and tested to enhance the consistency of the soft prompt tuning process. Through extensive experimentation with various combinations, we achieved an extraction rate of 65.2% at a 50-token prefix comparison. Our comparisons of CoSPED with other reference works confirm our superior extraction rates. We evaluate CoSPED on more scenarios, achieving Pythia model extraction rate of 51.7% and introducing cross-model comparison. Finally, we explore defense through Rank-One Model Editing and achieve a reduction in the extraction rate to 1.6%, which proves that our analysis of extraction mechanisms can directly inform effective mitigation strategies against soft prompt-based attacks.
title CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
topic Cryptography and Security
url https://arxiv.org/abs/2510.11137