The Great Pretender: A Stochasticity Problem in LLM Jailbreak
Fuente:
arXiv
Saved in:
| Main Authors: | Monteuuis, Jean-Philippe, Chen, Cong, Petit, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
The Echo Chamber Multi-Turn LLM Jailbreak
by: Alobaid, Ahmad, et al.
Published: (2026)
by: Alobaid, Ahmad, et al.
Published: (2026)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025)
by: Wu, ChenYu, et al.
Published: (2025)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
by: Zhang, Yingjie, et al.
Published: (2025)
by: Zhang, Yingjie, et al.
Published: (2025)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
by: Wang, Zhaoqi, et al.
Published: (2025)
by: Wang, Zhaoqi, et al.
Published: (2025)
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
by: Wang, Ziwei, et al.
Published: (2026)
by: Wang, Ziwei, et al.
Published: (2026)
LLM Jailbreak Detection for (Almost) Free!
by: Chen, Guorui, et al.
Published: (2025)
by: Chen, Guorui, et al.
Published: (2025)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
by: Tsmindashvili, Tatia, et al.
Published: (2025)
by: Tsmindashvili, Tatia, et al.
Published: (2025)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
by: Zhu, Wentian, et al.
Published: (2025)
by: Zhu, Wentian, et al.
Published: (2025)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
by: Qi, Senmao, et al.
Published: (2025)
by: Qi, Senmao, et al.
Published: (2025)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
by: Asl, Javad Rafiei, et al.
Published: (2025)
by: Asl, Javad Rafiei, et al.
Published: (2025)
CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
by: Hu, Jiaming, et al.
Published: (2025)
by: Hu, Jiaming, et al.
Published: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
by: Chen, Kejia, et al.
Published: (2026)
by: Chen, Kejia, et al.
Published: (2026)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
by: Bullwinkel, Blake, et al.
Published: (2025)
by: Bullwinkel, Blake, et al.
Published: (2025)
EVA: Editing for Versatile Alignment against Jailbreaks
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Emoji-Based Jailbreaking of Large Language Models
by: Gopinadh, M P V S, et al.
Published: (2026)
by: Gopinadh, M P V S, et al.
Published: (2026)
Jailbreaking is (Mostly) Simpler Than You Think
by: Russinovich, Mark, et al.
Published: (2025)
by: Russinovich, Mark, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
by: Xiong, Chen, et al.
Published: (2026)
by: Xiong, Chen, et al.
Published: (2026)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
by: Chu, Junjie, et al.
Published: (2025)
by: Chu, Junjie, et al.
Published: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
by: Li, Jie, et al.
Published: (2024)
by: Li, Jie, et al.
Published: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
by: Wang, Zi, et al.
Published: (2024)
by: Wang, Zi, et al.
Published: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
by: Pramanik, Vishal, et al.
Published: (2026)
by: Pramanik, Vishal, et al.
Published: (2026)
Similar Items
-
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
by: Russinovich, Mark, et al.
Published: (2024) -
The Echo Chamber Multi-Turn LLM Jailbreak
by: Alobaid, Ahmad, et al.
Published: (2026) -
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025) -
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
by: Zhang, Yingjie, et al.
Published: (2025) -
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)