Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
Fuente:
arXiv
Saved in:
| Main Authors: | Derya, Kemal, Sunar, Berk |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FAULT+PROBE: A Generic Rowhammer-based Bit Recovery Attack
by: Derya, Kemal, et al.
Published: (2024)
by: Derya, Kemal, et al.
Published: (2024)
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
by: Tol, M. Caner, et al.
Published: (2025)
by: Tol, M. Caner, et al.
Published: (2025)
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
LeapFrog: The Rowhammer Instruction Skip Attack
by: Adiletta, Andrew, et al.
Published: (2024)
by: Adiletta, Andrew, et al.
Published: (2024)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)
by: Zhang, Shenyi, et al.
Published: (2025)
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
Mayhem: Targeted Corruption of Register and Stack Variables
by: Adiletta, Andrew J., et al.
Published: (2023)
by: Adiletta, Andrew J., et al.
Published: (2023)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
TRYLOCK: Defense-in-Depth Against LLM Jailbreaks via Layered Preference and Representation Engineering
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
by: Mishra, Rina, et al.
Published: (2025)
by: Mishra, Rina, et al.
Published: (2025)
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
by: Chen, Luoyu, et al.
Published: (2026)
by: Chen, Luoyu, et al.
Published: (2026)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
by: Zeng, Xiyu, et al.
Published: (2025)
by: Zeng, Xiyu, et al.
Published: (2025)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
by: Mai, Wuyuao, et al.
Published: (2025)
by: Mai, Wuyuao, et al.
Published: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
by: Weng, Fenghua, et al.
Published: (2024)
by: Weng, Fenghua, et al.
Published: (2024)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
by: Yang, Zhuoran, et al.
Published: (2025)
by: Yang, Zhuoran, et al.
Published: (2025)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
by: Mao, Yanxu, et al.
Published: (2025)
by: Mao, Yanxu, et al.
Published: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
by: Cunningham, Hoagy, et al.
Published: (2026)
by: Cunningham, Hoagy, et al.
Published: (2026)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
by: Liang, Siyuan, et al.
Published: (2025)
by: Liang, Siyuan, et al.
Published: (2025)
A Learning-Based Attack Framework to Break SOTA Poisoning Defenses in Federated Learning
by: Yang, Yuxin, et al.
Published: (2024)
by: Yang, Yuxin, et al.
Published: (2024)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
by: Kim, Taeyoun, et al.
Published: (2024)
by: Kim, Taeyoun, et al.
Published: (2024)
DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Similar Items
-
FAULT+PROBE: A Generic Rowhammer-based Bit Recovery Attack
by: Derya, Kemal, et al.
Published: (2024) -
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
by: Tol, M. Caner, et al.
Published: (2025) -
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
by: Adiletta, Andrew, et al.
Published: (2025) -
LeapFrog: The Rowhammer Instruction Skip Attack
by: Adiletta, Andrew, et al.
Published: (2024) -
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)