How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Songyang, Li, Chaozhuo, Pu, Rui, Zhang, Litian, Wang, Chenxu, Chen, Zejian, Zhang, Yuting, Hei, Yiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
von: Pu, Rui, et al.
Veröffentlicht: (2024)
von: Pu, Rui, et al.
Veröffentlicht: (2024)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
von: Pu, Rui, et al.
Veröffentlicht: (2025)
von: Pu, Rui, et al.
Veröffentlicht: (2025)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2024)
von: Ran, Delong, et al.
Veröffentlicht: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
von: Wu, Yu-Hang, et al.
Veröffentlicht: (2025)
von: Wu, Yu-Hang, et al.
Veröffentlicht: (2025)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
Proactive defense against LLM Jailbreak
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Imperceptible Jailbreaking against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
von: Kadali, Sri Durga Sai Sowmya, et al.
Veröffentlicht: (2026)
von: Kadali, Sri Durga Sai Sowmya, et al.
Veröffentlicht: (2026)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
von: He, Zeqing, et al.
Veröffentlicht: (2024)
von: He, Zeqing, et al.
Veröffentlicht: (2024)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
von: Cheng, Siyang, et al.
Veröffentlicht: (2025)
von: Cheng, Siyang, et al.
Veröffentlicht: (2025)
Jailbreak Distillation: Renewable Safety Benchmarking
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Jailbreaking in the Haystack
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
von: Li, Jindong, et al.
Veröffentlicht: (2026)
von: Li, Jindong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
von: Chen, Zejian, et al.
Veröffentlicht: (2026) -
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
von: Pu, Rui, et al.
Veröffentlicht: (2024) -
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
von: Liu, Songyang, et al.
Veröffentlicht: (2025) -
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
von: Pu, Rui, et al.
Veröffentlicht: (2025) -
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
von: Liu, Songyang, et al.
Veröffentlicht: (2026)