JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xiaoyu, Zhang, Cen, Li, Tianlin, Huang, Yihao, Jia, Xiaojun, Hu, Ming, Zhang, Jie, Liu, Yang, Ma, Shiqing, Shen, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation
by: Zhang, Xiaolei, et al.
Published: (2026)
by: Zhang, Xiaolei, et al.
Published: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
by: Zhang, Xinkai, et al.
Published: (2026)
by: Zhang, Xinkai, et al.
Published: (2026)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
by: Zhao, Wei, et al.
Published: (2026)
by: Zhao, Wei, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
by: Luo, Jiaqi, et al.
Published: (2026)
by: Luo, Jiaqi, et al.
Published: (2026)
MirGuard: Towards a Robust Provenance-based Intrusion Detection System Against Graph Manipulation Attacks
by: Sang, Anyuan, et al.
Published: (2025)
by: Sang, Anyuan, et al.
Published: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
by: Luo, Weidi, et al.
Published: (2024)
by: Luo, Weidi, et al.
Published: (2024)
Sealing the Audit-Runtime Gap for LLM Skills
by: Shen, Tingda, et al.
Published: (2026)
by: Shen, Tingda, et al.
Published: (2026)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
by: Yeke, Doguhuan, et al.
Published: (2026)
by: Yeke, Doguhuan, et al.
Published: (2026)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)
by: Jiang, Weipeng, et al.
Published: (2026)
SyncGuard: Robust Audio Watermarking Capable of Countering Desynchronization Attacks
by: Gan, Zhenliang, et al.
Published: (2025)
by: Gan, Zhenliang, et al.
Published: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
by: Zou, Wei, et al.
Published: (2025)
by: Zou, Wei, et al.
Published: (2025)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
by: Teng, Ma, et al.
Published: (2024)
by: Teng, Ma, et al.
Published: (2024)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2026)
by: Jia, Yuqi, et al.
Published: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
by: Wang, Zhilong, et al.
Published: (2025)
by: Wang, Zhilong, et al.
Published: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
by: Zhang, Wenhui, et al.
Published: (2026)
by: Zhang, Wenhui, et al.
Published: (2026)
ClawGuard: Out-of-Band Detection of LLM Agent Workflow Hijacking via EM Side Channel
by: Gan, Leo Linqian, et al.
Published: (2026)
by: Gan, Leo Linqian, et al.
Published: (2026)
Efficient and Universal Watermarking for LLM-Generated Code Detection
by: Li, Boquan, et al.
Published: (2024)
by: Li, Boquan, et al.
Published: (2024)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
by: Chen, Jing, et al.
Published: (2025)
by: Chen, Jing, et al.
Published: (2025)
Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference Attacks
by: Noorbakhsh, Sayedeh Leila, et al.
Published: (2024)
by: Noorbakhsh, Sayedeh Leila, et al.
Published: (2024)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
by: Yin, Yu, et al.
Published: (2026)
by: Yin, Yu, et al.
Published: (2026)
Following Devils' Footprint: Towards Real-time Detection of Price Manipulation Attacks
by: Zhang, Bosi, et al.
Published: (2025)
by: Zhang, Bosi, et al.
Published: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
by: Wang, Reachal, et al.
Published: (2025)
by: Wang, Reachal, et al.
Published: (2025)
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
by: Koide, Takashi, et al.
Published: (2026)
by: Koide, Takashi, et al.
Published: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
CipherGuard: Compiler-aided Mitigation against Ciphertext Side-channel Attacks
by: Jiang, Ke, et al.
Published: (2025)
by: Jiang, Ke, et al.
Published: (2025)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)
by: Liu, Zesen, et al.
Published: (2025)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
Similar Items
-
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026) -
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
by: Wang, Xi, et al.
Published: (2025) -
CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation
by: Zhang, Xiaolei, et al.
Published: (2026) -
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
by: Zhang, Xinkai, et al.
Published: (2026) -
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)