Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
Fuente:
arXiv
Guardado en:
| Autores principales: | Russinovich, Mark, Salem, Ahmed, Eldan, Ronen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jailbreaking is (Mostly) Simpler Than You Think
por: Russinovich, Mark, et al.
Publicado: (2025)
por: Russinovich, Mark, et al.
Publicado: (2025)
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
por: Russinovich, Mark, et al.
Publicado: (2024)
por: Russinovich, Mark, et al.
Publicado: (2024)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
por: Bullwinkel, Blake, et al.
Publicado: (2025)
por: Bullwinkel, Blake, et al.
Publicado: (2025)
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
por: Russinovich, Mark, et al.
Publicado: (2025)
por: Russinovich, Mark, et al.
Publicado: (2025)
The Echo Chamber Multi-Turn LLM Jailbreak
por: Alobaid, Ahmad, et al.
Publicado: (2026)
por: Alobaid, Ahmad, et al.
Publicado: (2026)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
por: Lin, Xingwei, et al.
Publicado: (2026)
por: Lin, Xingwei, et al.
Publicado: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
por: Zhang, Xinkai, et al.
Publicado: (2026)
por: Zhang, Xinkai, et al.
Publicado: (2026)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
por: Wu, ChenYu, et al.
Publicado: (2025)
por: Wu, ChenYu, et al.
Publicado: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
por: Gibbs, Tom, et al.
Publicado: (2024)
por: Gibbs, Tom, et al.
Publicado: (2024)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
por: Narula, Sidhant, et al.
Publicado: (2025)
por: Narula, Sidhant, et al.
Publicado: (2025)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
por: Asl, Javad Rafiei, et al.
Publicado: (2025)
por: Asl, Javad Rafiei, et al.
Publicado: (2025)
The Great Pretender: A Stochasticity Problem in LLM Jailbreak
por: Monteuuis, Jean-Philippe, et al.
Publicado: (2026)
por: Monteuuis, Jean-Philippe, et al.
Publicado: (2026)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
por: Qi, Senmao, et al.
Publicado: (2025)
por: Qi, Senmao, et al.
Publicado: (2025)
Securing AI Agents with Information-Flow Control
por: Costa, Manuel, et al.
Publicado: (2025)
por: Costa, Manuel, et al.
Publicado: (2025)
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
por: Kulkarni, Prashant
Publicado: (2026)
por: Kulkarni, Prashant
Publicado: (2026)
Untargeted Jailbreak Attack
por: Huang, Xinzhe, et al.
Publicado: (2025)
por: Huang, Xinzhe, et al.
Publicado: (2025)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
por: Tsmindashvili, Tatia, et al.
Publicado: (2025)
por: Tsmindashvili, Tatia, et al.
Publicado: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
por: Das, Badhan Chandra, et al.
Publicado: (2026)
por: Das, Badhan Chandra, et al.
Publicado: (2026)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
por: Zhang, Yingjie, et al.
Publicado: (2025)
por: Zhang, Yingjie, et al.
Publicado: (2025)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)
por: Aqrawi, Alan, et al.
Publicado: (2024)
por: Aqrawi, Alan, et al.
Publicado: (2024)
Involuntary Jailbreak: On Self-Prompting Attacks
por: Guo, Yangyang, et al.
Publicado: (2025)
por: Guo, Yangyang, et al.
Publicado: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
por: Hossain, Ismail, et al.
Publicado: (2026)
por: Hossain, Ismail, et al.
Publicado: (2026)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
por: Jiang, Shuli, et al.
Publicado: (2024)
por: Jiang, Shuli, et al.
Publicado: (2024)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)
por: Wen, Rui, et al.
Publicado: (2026)
Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks
por: You, Doohee
Publicado: (2026)
por: You, Doohee
Publicado: (2026)
FlipAttack: Jailbreak LLMs via Flipping
por: Liu, Yue, et al.
Publicado: (2024)
por: Liu, Yue, et al.
Publicado: (2024)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
por: Hong, Wenjing, et al.
Publicado: (2026)
por: Hong, Wenjing, et al.
Publicado: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
por: Ma, Jiachen, et al.
Publicado: (2024)
por: Ma, Jiachen, et al.
Publicado: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
por: Jaiswal, Piyush, et al.
Publicado: (2026)
por: Jaiswal, Piyush, et al.
Publicado: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
por: Zhang, Zheng, et al.
Publicado: (2025)
por: Zhang, Zheng, et al.
Publicado: (2025)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
por: Wang, Zhilong, et al.
Publicado: (2024)
por: Wang, Zhilong, et al.
Publicado: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
por: Yu, Miao, et al.
Publicado: (2024)
por: Yu, Miao, et al.
Publicado: (2024)
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
por: Zhou, Andy, et al.
Publicado: (2025)
por: Zhou, Andy, et al.
Publicado: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
por: Park, Junyoung, et al.
Publicado: (2026)
por: Park, Junyoung, et al.
Publicado: (2026)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
por: Shang, Zhengchun, et al.
Publicado: (2025)
por: Shang, Zhengchun, et al.
Publicado: (2025)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
por: Yoon, Sangyeon, et al.
Publicado: (2026)
por: Yoon, Sangyeon, et al.
Publicado: (2026)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
por: Chen, Kejia, et al.
Publicado: (2026)
por: Chen, Kejia, et al.
Publicado: (2026)
Ejemplares similares
-
Jailbreaking is (Mostly) Simpler Than You Think
por: Russinovich, Mark, et al.
Publicado: (2025) -
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
por: Russinovich, Mark, et al.
Publicado: (2024) -
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
por: Bullwinkel, Blake, et al.
Publicado: (2025) -
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
por: Russinovich, Mark, et al.
Publicado: (2025) -
The Echo Chamber Multi-Turn LLM Jailbreak
por: Alobaid, Ahmad, et al.
Publicado: (2026)