SoK: Evaluating Jailbreak Guardrails for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xunguang, Ji, Zhenlan, Wang, Wenxuan, Li, Zongjie, Wu, Daoyuan, Wang, Shuai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
SoK: Semantic Privacy in Large Language Models
by: Ma, Baihe, et al.
Published: (2025)
by: Ma, Baihe, et al.
Published: (2025)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
by: Hong, Hanbin, et al.
Published: (2025)
by: Hong, Hanbin, et al.
Published: (2025)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
SoK: On Gradient Leakage in Federated Learning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
SoK: Prompt Hacking of Large Language Models
by: Rababah, Baha, et al.
Published: (2024)
by: Rababah, Baha, et al.
Published: (2024)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
by: Shao, Shuo, et al.
Published: (2025)
by: Shao, Shuo, et al.
Published: (2025)
SoK: Unlearnability and Unlearning for Model Dememorization
by: Zhang, Mengying, et al.
Published: (2026)
by: Zhang, Mengying, et al.
Published: (2026)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
by: Abad, Gorka, et al.
Published: (2025)
by: Abad, Gorka, et al.
Published: (2025)
SoK: Trust-Authorization Mismatch in LLM Agent Interactions
by: Shi, Guanquan, et al.
Published: (2025)
by: Shi, Guanquan, et al.
Published: (2025)
SoK: On the Semantic AI Security in Autonomous Driving
by: Shen, Junjie, et al.
Published: (2022)
by: Shen, Junjie, et al.
Published: (2022)
SoK: Security and Privacy of AI Agents for Blockchain
by: Romandini, Nicolò, et al.
Published: (2025)
by: Romandini, Nicolò, et al.
Published: (2025)
SoK: Towards Security and Safety of Edge AI
by: Wingarz, Tatjana, et al.
Published: (2024)
by: Wingarz, Tatjana, et al.
Published: (2024)
SoK: Watermarking for AI-Generated Content
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025)
by: Wu, ChenYu, et al.
Published: (2025)
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
SEAL: Subspace-Anchored Watermarks for LLM Ownership
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
by: Wu, Qilong, et al.
Published: (2025)
by: Wu, Qilong, et al.
Published: (2025)
SoK: Blockchain-Based Decentralized AI (DeAI)
by: Lui, Elizabeth, et al.
Published: (2024)
by: Lui, Elizabeth, et al.
Published: (2024)
SoK: How Robust is Audio Watermarking in Generative AI models?
by: Wen, Yizhu, et al.
Published: (2025)
by: Wen, Yizhu, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
SoK: Verifiable Cross-Silo FL
by: Korneev, Aleksei, et al.
Published: (2024)
by: Korneev, Aleksei, et al.
Published: (2024)
SoK: On the Offensive Potential of AI
by: Schröer, Saskia Laura, et al.
Published: (2024)
by: Schröer, Saskia Laura, et al.
Published: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
SoK: Security and Privacy Risks of Healthcare AI
by: Chang, Yuanhaur, et al.
Published: (2024)
by: Chang, Yuanhaur, et al.
Published: (2024)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
by: Wang, Shenao, et al.
Published: (2025)
by: Wang, Shenao, et al.
Published: (2025)
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
by: Wu, Zihui, et al.
Published: (2024)
by: Wu, Zihui, et al.
Published: (2024)
SoK: Enhancing Cryptographic Collaborative Learning with Differential Privacy
by: Capano, Francesco, et al.
Published: (2026)
by: Capano, Francesco, et al.
Published: (2026)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Similar Items
-
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025) -
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024) -
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026) -
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025) -
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)