The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Zhili, Xi, Zihang, He, Ying, Tong, Wei, Hua, Jingyu, Zhong, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Balancing Innovation and Privacy: Data Security Strategies in Natural Language Processing Applications
by: Liu, Shaobo, et al.
Published: (2024)
by: Liu, Shaobo, et al.
Published: (2024)
Reverse-Engineering Model Editing on Language Models
by: Sun, Zhiyu, et al.
Published: (2026)
by: Sun, Zhiyu, et al.
Published: (2026)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
by: Sun, Xiongtao, et al.
Published: (2024)
by: Sun, Xiongtao, et al.
Published: (2024)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
by: Li, Zi, et al.
Published: (2026)
by: Li, Zi, et al.
Published: (2026)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
by: Agarwal, Divyansh, et al.
Published: (2024)
by: Agarwal, Divyansh, et al.
Published: (2024)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
by: Lin, Sam, et al.
Published: (2024)
by: Lin, Sam, et al.
Published: (2024)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
A High-Capacity and Secure Disambiguation Algorithm for Neural Linguistic Steganography
by: Feng, Yapei, et al.
Published: (2025)
by: Feng, Yapei, et al.
Published: (2025)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
by: Zeng, Shenglai, et al.
Published: (2024)
by: Zeng, Shenglai, et al.
Published: (2024)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
by: Guo, Ruoqi, et al.
Published: (2026)
by: Guo, Ruoqi, et al.
Published: (2026)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
by: Ostermann, Simon, et al.
Published: (2024)
by: Ostermann, Simon, et al.
Published: (2024)
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
by: Gu, Chenxi, et al.
Published: (2026)
by: Gu, Chenxi, et al.
Published: (2026)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
by: Das, Saswat, et al.
Published: (2026)
by: Das, Saswat, et al.
Published: (2026)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
Security and Privacy Challenges of Large Language Models: A Survey
by: Das, Badhan Chandra, et al.
Published: (2024)
by: Das, Badhan Chandra, et al.
Published: (2024)
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
by: Xiao, Yunze, et al.
Published: (2026)
by: Xiao, Yunze, et al.
Published: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
by: Pham, Viet, et al.
Published: (2025)
by: Pham, Viet, et al.
Published: (2025)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
by: Mohammadi, Bardia, et al.
Published: (2026)
by: Mohammadi, Bardia, et al.
Published: (2026)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
by: Li, Jianwei, et al.
Published: (2023)
by: Li, Jianwei, et al.
Published: (2023)
Optimizing the Privacy-Utility Balance using Synthetic Data and Configurable Perturbation Pipelines
by: Sharma, Anantha, et al.
Published: (2025)
by: Sharma, Anantha, et al.
Published: (2025)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024)
by: Qiu, Huming, et al.
Published: (2024)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
by: Luo, Zhifan, et al.
Published: (2025)
by: Luo, Zhifan, et al.
Published: (2025)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
by: Shen, Hao, et al.
Published: (2025)
by: Shen, Hao, et al.
Published: (2025)
Similar Items
-
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
by: Zhang, Tianyu, et al.
Published: (2025) -
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024) -
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024) -
Balancing Innovation and Privacy: Data Security Strategies in Natural Language Processing Applications
by: Liu, Shaobo, et al.
Published: (2024) -
Reverse-Engineering Model Editing on Language Models
by: Sun, Zhiyu, et al.
Published: (2026)