Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
Fuente:
arXiv
Guardado en:
| Autores principales: | Kang, Daewon, Shin, YeongHwan, Kim, Doyeon, Jung, Kyu-Hwan, Son, Meong Hi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025)
por: Shi, Jiawen, et al.
Publicado: (2025)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
por: Wang, Yihan, et al.
Publicado: (2025)
por: Wang, Yihan, et al.
Publicado: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
por: Hossain, S M Asif, et al.
Publicado: (2025)
por: Hossain, S M Asif, et al.
Publicado: (2025)
Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
por: Yang, Yulong, et al.
Publicado: (2023)
por: Yang, Yulong, et al.
Publicado: (2023)
Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses
por: Song, Minkyoo, et al.
Publicado: (2026)
por: Song, Minkyoo, et al.
Publicado: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025)
por: Maloyan, Narek, et al.
Publicado: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
por: Yin, Yu, et al.
Publicado: (2026)
por: Yin, Yu, et al.
Publicado: (2026)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
por: Wang, Reachal, et al.
Publicado: (2025)
por: Wang, Reachal, et al.
Publicado: (2025)
Poison Attacks and Adversarial Prompts Against an Informed University Virtual Assistant
por: Fernandez, Ivan A., et al.
Publicado: (2024)
por: Fernandez, Ivan A., et al.
Publicado: (2024)
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
por: Du, Chenghao, et al.
Publicado: (2025)
por: Du, Chenghao, et al.
Publicado: (2025)
Adversarial Attacks on Reinforcement Learning Agents for Command and Control
por: Dabholkar, Ahaan, et al.
Publicado: (2024)
por: Dabholkar, Ahaan, et al.
Publicado: (2024)
Explainable and Transferable Adversarial Attack for ML-Based Network Intrusion Detectors
por: Zhang, Hangsheng, et al.
Publicado: (2024)
por: Zhang, Hangsheng, et al.
Publicado: (2024)
PINA: Prompt Injection Attack against Navigation Agents
por: Liu, Jiani, et al.
Publicado: (2026)
por: Liu, Jiani, et al.
Publicado: (2026)
HETAL: Efficient Privacy-preserving Transfer Learning with Homomorphic Encryption
por: Lee, Seewoo, et al.
Publicado: (2024)
por: Lee, Seewoo, et al.
Publicado: (2024)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
por: Liu, Zesen, et al.
Publicado: (2025)
por: Liu, Zesen, et al.
Publicado: (2025)
Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
por: Chang, Xiangyu, et al.
Publicado: (2025)
por: Chang, Xiangyu, et al.
Publicado: (2025)
Exploring the Robustness and Transferability of Patch-Based Adversarial Attacks in Quantized Neural Networks
por: Guesmi, Amira, et al.
Publicado: (2024)
por: Guesmi, Amira, et al.
Publicado: (2024)
Toward Realistic Adversarial Attacks in IDS: A Novel Feasibility Metric for Transferability
por: Ennaji, Sabrine, et al.
Publicado: (2025)
por: Ennaji, Sabrine, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
por: Ma, Jiachen, et al.
Publicado: (2024)
por: Ma, Jiachen, et al.
Publicado: (2024)
The Relationship Between Network Similarity and Transferability of Adversarial Attacks
por: Klause, Gerrit, et al.
Publicado: (2025)
por: Klause, Gerrit, et al.
Publicado: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025)
por: Zizzo, Giulio, et al.
Publicado: (2025)
MalTool: Malicious Tool Attacks on LLM Agents
por: Hu, Yuepeng, et al.
Publicado: (2026)
por: Hu, Yuepeng, et al.
Publicado: (2026)
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
por: Pandey, Rohan, et al.
Publicado: (2026)
por: Pandey, Rohan, et al.
Publicado: (2026)
Adversarial Contrastive Learning for LLM Quantization Attacks
por: Song, Dinghong, et al.
Publicado: (2026)
por: Song, Dinghong, et al.
Publicado: (2026)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
por: Kang, Mintong, et al.
Publicado: (2023)
por: Kang, Mintong, et al.
Publicado: (2023)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
Mutual Information Minimization for Side-Channel Attack Resistance via Optimal Noise Injection
por: Woo, Jiheon, et al.
Publicado: (2025)
por: Woo, Jiheon, et al.
Publicado: (2025)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
por: Johnson, Sam, et al.
Publicado: (2025)
por: Johnson, Sam, et al.
Publicado: (2025)
Enhancing TinyML Security: Study of Adversarial Attack Transferability
por: Shah, Parin, et al.
Publicado: (2024)
por: Shah, Parin, et al.
Publicado: (2024)
Enhancing Adversarial Transferability with Adversarial Weight Tuning
por: Chen, Jiahao, et al.
Publicado: (2024)
por: Chen, Jiahao, et al.
Publicado: (2024)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
por: Hong, Hanbin, et al.
Publicado: (2023)
por: Hong, Hanbin, et al.
Publicado: (2023)
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
por: Koide, Takashi, et al.
Publicado: (2026)
por: Koide, Takashi, et al.
Publicado: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
por: Chen, Yulin, et al.
Publicado: (2026)
por: Chen, Yulin, et al.
Publicado: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
por: Zhuang, Zhixiong, et al.
Publicado: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
Ejemplares similares
-
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025) -
Prompt Injection Attack to Tool Selection in LLM Agents
por: Shi, Jiawen, et al.
Publicado: (2025) -
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
por: Wang, Yihan, et al.
Publicado: (2025) -
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
por: Hossain, S M Asif, et al.
Publicado: (2025) -
Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
por: Yang, Yulong, et al.
Publicado: (2023)