DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Xiongtao, Liu, Gan, He, Zhipeng, Li, Hui, Li, Xiaoguang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
di: Zheng, Yujia, et al.
Pubblicazione: (2025)
di: Zheng, Yujia, et al.
Pubblicazione: (2025)
Efficient Detection of Toxic Prompts in Large Language Models
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
Goal-guided Generative Prompt Injection Attack on Large Language Models
di: Zhang, Chong, et al.
Pubblicazione: (2024)
di: Zhang, Chong, et al.
Pubblicazione: (2024)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
di: Gong, Yichen, et al.
Pubblicazione: (2023)
di: Gong, Yichen, et al.
Pubblicazione: (2023)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Evaluation of Prompt Injection Defenses in Large Language Models
di: Deep, Priyal, et al.
Pubblicazione: (2026)
di: Deep, Priyal, et al.
Pubblicazione: (2026)
SoK: Prompt Hacking of Large Language Models
di: Rababah, Baha, et al.
Pubblicazione: (2024)
di: Rababah, Baha, et al.
Pubblicazione: (2024)
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
di: Chen, Baicheng, et al.
Pubblicazione: (2026)
di: Chen, Baicheng, et al.
Pubblicazione: (2026)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
di: Hong, Hanbin, et al.
Pubblicazione: (2025)
di: Hong, Hanbin, et al.
Pubblicazione: (2025)
An Engorgio Prompt Makes Large Language Model Babble on
di: Dong, Jianshuo, et al.
Pubblicazione: (2024)
di: Dong, Jianshuo, et al.
Pubblicazione: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
di: Wang, Libo
Pubblicazione: (2024)
di: Wang, Libo
Pubblicazione: (2024)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
di: Ganiuly, Daniyal, et al.
Pubblicazione: (2025)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
di: Chugh, Rishit
Pubblicazione: (2026)
di: Chugh, Rishit
Pubblicazione: (2026)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
di: Levin, Roman, et al.
Pubblicazione: (2025)
di: Levin, Roman, et al.
Pubblicazione: (2025)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
di: Li, Zongze, et al.
Pubblicazione: (2025)
di: Li, Zongze, et al.
Pubblicazione: (2025)
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
di: Shen, Zhili, et al.
Pubblicazione: (2024)
di: Shen, Zhili, et al.
Pubblicazione: (2024)
Prompt Injection as Role Confusion
di: Ye, Charles, et al.
Pubblicazione: (2026)
di: Ye, Charles, et al.
Pubblicazione: (2026)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
di: Ran, Delong, et al.
Pubblicazione: (2024)
di: Ran, Delong, et al.
Pubblicazione: (2024)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
di: Xiong, Junjie, et al.
Pubblicazione: (2025)
di: Xiong, Junjie, et al.
Pubblicazione: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
di: Li, Xirui, et al.
Pubblicazione: (2024)
di: Li, Xirui, et al.
Pubblicazione: (2024)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
di: Lv, Huijie, et al.
Pubblicazione: (2024)
di: Lv, Huijie, et al.
Pubblicazione: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Resource Consumption Threats in Large Language Models
di: Zhang, Yuanhe, et al.
Pubblicazione: (2026)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
Have You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging
di: Cong, Tianshuo, et al.
Pubblicazione: (2024)
di: Cong, Tianshuo, et al.
Pubblicazione: (2024)
Code Vulnerability Repair with Large Language Model using Context-Aware Prompt Tuning
di: Khan, Arshiya, et al.
Pubblicazione: (2024)
di: Khan, Arshiya, et al.
Pubblicazione: (2024)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
di: Cao, Bochuan, et al.
Pubblicazione: (2025)
di: Cao, Bochuan, et al.
Pubblicazione: (2025)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
PIArena: A Platform for Prompt Injection Evaluation
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
PromptKeeper: Safeguarding System Prompts for LLMs
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
di: Shao, Shuo, et al.
Pubblicazione: (2025)
di: Shao, Shuo, et al.
Pubblicazione: (2025)
Internal Safety Collapse in Frontier Large Language Models
di: Wu, Yutao, et al.
Pubblicazione: (2026)
di: Wu, Yutao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
di: Zheng, Yujia, et al.
Pubblicazione: (2025) -
Efficient Detection of Toxic Prompts in Large Language Models
di: Liu, Yi, et al.
Pubblicazione: (2024) -
Goal-guided Generative Prompt Injection Attack on Large Language Models
di: Zhang, Chong, et al.
Pubblicazione: (2024) -
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
di: Gong, Yichen, et al.
Pubblicazione: (2023) -
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)