Prompt Leakage effect and defense strategies for multi-turn LLM interactions
Fuente:
arXiv
Guardado en:
| Autores principales: | Agarwal, Divyansh, Fabbri, Alexander R., Risher, Ben, Laban, Philippe, Joty, Shafiq, Wu, Chien-Sheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Proactive defense against LLM Jailbreak
por: Zhao, Weiliang, et al.
Publicado: (2025)
por: Zhao, Weiliang, et al.
Publicado: (2025)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
por: Kumarage, Tharindu, et al.
Publicado: (2025)
por: Kumarage, Tharindu, et al.
Publicado: (2025)
Nonmalleable Progress Leakage
por: Cecchetti, Ethan
Publicado: (2025)
por: Cecchetti, Ethan
Publicado: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
por: Wang, Liwen, et al.
Publicado: (2025)
por: Wang, Liwen, et al.
Publicado: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Exploiting Leakage in Password Managers via Injection Attacks
por: Fábrega, Andrés, et al.
Publicado: (2024)
por: Fábrega, Andrés, et al.
Publicado: (2024)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
por: Cao, Bochuan, et al.
Publicado: (2025)
por: Cao, Bochuan, et al.
Publicado: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
por: Zhong, Peter Yong, et al.
Publicado: (2025)
por: Zhong, Peter Yong, et al.
Publicado: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
por: Freenor, Michael, et al.
Publicado: (2025)
por: Freenor, Michael, et al.
Publicado: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
por: Li, Caihua, et al.
Publicado: (2024)
por: Li, Caihua, et al.
Publicado: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
por: Choi, Jacob, et al.
Publicado: (2025)
por: Choi, Jacob, et al.
Publicado: (2025)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
por: Hu, Hanjiang, et al.
Publicado: (2025)
por: Hu, Hanjiang, et al.
Publicado: (2025)
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
por: Sekar, Anirudh, et al.
Publicado: (2026)
por: Sekar, Anirudh, et al.
Publicado: (2026)
Certifying LLM Safety against Adversarial Prompting
por: Kumar, Aounon, et al.
Publicado: (2023)
por: Kumar, Aounon, et al.
Publicado: (2023)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
por: Wang, Fei, et al.
Publicado: (2025)
por: Wang, Fei, et al.
Publicado: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025)
por: Maloyan, Narek, et al.
Publicado: (2025)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
por: Jawad, Huseein, et al.
Publicado: (2025)
por: Jawad, Huseein, et al.
Publicado: (2025)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
por: Noah, Amit Finkman, et al.
Publicado: (2024)
por: Noah, Amit Finkman, et al.
Publicado: (2024)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
por: Xin, Yuan, et al.
Publicado: (2026)
por: Xin, Yuan, et al.
Publicado: (2026)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
por: Singh, Himanshu, et al.
Publicado: (2026)
por: Singh, Himanshu, et al.
Publicado: (2026)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
por: Hughes, Anthony, et al.
Publicado: (2025)
por: Hughes, Anthony, et al.
Publicado: (2025)
Active Sybil attack and efficient defense strategy in IPFS DHT
por: Netto, V. H. de Moura, et al.
Publicado: (2025)
por: Netto, V. H. de Moura, et al.
Publicado: (2025)
BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
por: Gan, Jialing, et al.
Publicado: (2026)
por: Gan, Jialing, et al.
Publicado: (2026)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
por: Roy, Soham, et al.
Publicado: (2026)
por: Roy, Soham, et al.
Publicado: (2026)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
por: Yang, Yong, et al.
Publicado: (2024)
por: Yang, Yong, et al.
Publicado: (2024)
TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code
por: Sternfeld, Alexander, et al.
Publicado: (2025)
por: Sternfeld, Alexander, et al.
Publicado: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
LeakDojo: Decoding the Leakage Threats of RAG Systems
por: Zhang, Maosen, et al.
Publicado: (2026)
por: Zhang, Maosen, et al.
Publicado: (2026)
Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection
por: Rahman, Md Abdur, et al.
Publicado: (2024)
por: Rahman, Md Abdur, et al.
Publicado: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
por: Liang, Zi, et al.
Publicado: (2024)
por: Liang, Zi, et al.
Publicado: (2024)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
por: Sternfeld, Alexander, et al.
Publicado: (2026)
por: Sternfeld, Alexander, et al.
Publicado: (2026)
Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines
por: Huang, Tao, et al.
Publicado: (2026)
por: Huang, Tao, et al.
Publicado: (2026)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
Fingerprinting LLMs via Prompt Injection
por: Hu, Yuepeng, et al.
Publicado: (2025)
por: Hu, Yuepeng, et al.
Publicado: (2025)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
por: Liu, Jinbo, et al.
Publicado: (2025)
por: Liu, Jinbo, et al.
Publicado: (2025)
Ejemplares similares
-
Proactive defense against LLM Jailbreak
por: Zhao, Weiliang, et al.
Publicado: (2025) -
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
por: Kumarage, Tharindu, et al.
Publicado: (2025) -
Nonmalleable Progress Leakage
por: Cecchetti, Ethan
Publicado: (2025) -
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
por: Wang, Liwen, et al.
Publicado: (2025) -
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)