You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Bochuan, Li, Changjiang, Cao, Yuanpu, Ge, Yameng, Wang, Ting, Chen, Jinghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918148706402304
author Cao, Bochuan
Li, Changjiang
Cao, Yuanpu
Ge, Yameng
Wang, Ting
Chen, Jinghui
author_facet Cao, Bochuan
Li, Changjiang
Cao, Yuanpu
Ge, Yameng
Wang, Ting
Chen, Jinghui
contents Large language models (LLMs) have been widely adopted across various applications, leveraging customized system prompts for diverse tasks. Facing potential system prompt leakage risks, model developers have implemented strategies to prevent leakage, primarily by disabling LLMs from repeating their context when encountering known attack patterns. However, it remains vulnerable to new and unforeseen prompt-leaking techniques. In this paper, we first introduce a simple yet effective prompt leaking attack to reveal such risks. Our attack is capable of extracting system prompts from various LLM-based application, even from SOTA LLM models such as GPT-4o or Claude 3.5 Sonnet. Our findings further inspire us to search for a fundamental solution to the problems by having no system prompt in the context. To this end, we propose SysVec, a novel method that encodes system prompts as internal representation vectors rather than raw text. By doing so, SysVec minimizes the risk of unauthorized disclosure while preserving the LLM's core language capabilities. Remarkably, this approach not only enhances security but also improves the model's general instruction-following abilities. Experimental results demonstrate that SysVec effectively mitigates prompt leakage attacks, preserves the LLM's functional integrity, and helps alleviate the forgetting issue in long-context scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
Cao, Bochuan
Li, Changjiang
Cao, Yuanpu
Ge, Yameng
Wang, Ting
Chen, Jinghui
Cryptography and Security
Artificial Intelligence
Computation and Language
Large language models (LLMs) have been widely adopted across various applications, leveraging customized system prompts for diverse tasks. Facing potential system prompt leakage risks, model developers have implemented strategies to prevent leakage, primarily by disabling LLMs from repeating their context when encountering known attack patterns. However, it remains vulnerable to new and unforeseen prompt-leaking techniques. In this paper, we first introduce a simple yet effective prompt leaking attack to reveal such risks. Our attack is capable of extracting system prompts from various LLM-based application, even from SOTA LLM models such as GPT-4o or Claude 3.5 Sonnet. Our findings further inspire us to search for a fundamental solution to the problems by having no system prompt in the context. To this end, we propose SysVec, a novel method that encodes system prompts as internal representation vectors rather than raw text. By doing so, SysVec minimizes the risk of unauthorized disclosure while preserving the LLM's core language capabilities. Remarkably, this approach not only enhances security but also improves the model's general instruction-following abilities. Experimental results demonstrate that SysVec effectively mitigates prompt leakage attacks, preserves the LLM's functional integrity, and helps alleviate the forgetting issue in long-context scenarios.
title You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.21884