Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sternak, Tvrtko, Runje, Davor, Granoša, Dorian, Wang, Chi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024)
von: Cui, Jing, et al.
Veröffentlicht: (2024)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
von: Li, Zongze, et al.
Veröffentlicht: (2025)
von: Li, Zongze, et al.
Veröffentlicht: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
von: Clusmann, Jan, et al.
Veröffentlicht: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks
von: Syed, Toqeer Ali, et al.
Veröffentlicht: (2025)
von: Syed, Toqeer Ali, et al.
Veröffentlicht: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
Network-Level Prompt and Trait Leakage in Local Research Agents
von: Jeong, Hyejun, et al.
Veröffentlicht: (2025)
von: Jeong, Hyejun, et al.
Veröffentlicht: (2025)
An Engorgio Prompt Makes Large Language Model Babble on
von: Dong, Jianshuo, et al.
Veröffentlicht: (2024)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2024)
A Survey of Attacks on Large Language Models
von: Xu, Wenrui, et al.
Veröffentlicht: (2025)
von: Xu, Wenrui, et al.
Veröffentlicht: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
von: Hui, Bo, et al.
Veröffentlicht: (2024)
von: Hui, Bo, et al.
Veröffentlicht: (2024)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Automating Security Audit Using Large Language Model based Agent: An Exploration Experiment
von: Chin, Jia Hui, et al.
Veröffentlicht: (2025)
von: Chin, Jia Hui, et al.
Veröffentlicht: (2025)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
Membership Inference Attacks on Tokenizers of Large Language Models
von: Tong, Meng, et al.
Veröffentlicht: (2025)
von: Tong, Meng, et al.
Veröffentlicht: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
von: Suo, Xuchen
Veröffentlicht: (2024)
von: Suo, Xuchen
Veröffentlicht: (2024)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
von: Zhao, Jin, et al.
Veröffentlicht: (2026)
von: Zhao, Jin, et al.
Veröffentlicht: (2026)
Unseen Attack Detection in Software-Defined Networking Using a BERT-Based Large Language Model
von: Swileh, Mohammed N., et al.
Veröffentlicht: (2024)
von: Swileh, Mohammed N., et al.
Veröffentlicht: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
TxRay: Agentic Postmortem of Live Blockchain Attacks
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
von: Levin, Roman, et al.
Veröffentlicht: (2025)
von: Levin, Roman, et al.
Veröffentlicht: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
AttacKG+:Boosting Attack Knowledge Graph Construction with Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
von: Hong, Hanbin, et al.
Veröffentlicht: (2025)
von: Hong, Hanbin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
von: Jiao, Yang, et al.
Veröffentlicht: (2025) -
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
von: Hong, Wenjing, et al.
Veröffentlicht: (2026) -
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
von: Wang, Che, et al.
Veröffentlicht: (2026) -
Recent Advances in Attack and Defense Approaches of Large Language Models
von: Cui, Jing, et al.
Veröffentlicht: (2024) -
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
von: Li, Zongze, et al.
Veröffentlicht: (2025)