The Echo Chamber Multi-Turn LLM Jailbreak
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alobaid, Ahmad, Roca, Martí Jordà, Castillo, Carlos, Vendrell, Joan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Introducing the Generative Application Firewall (GAF)
von: Farreny, Joan Vendrell, et al.
Veröffentlicht: (2026)
von: Farreny, Joan Vendrell, et al.
Veröffentlicht: (2026)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2025)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2025)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
Shattering the Echo Chamber: Hidden Safeguards in Manuscripts Against the AI Takeover of Peer Review
von: Ma, Oubo, et al.
Veröffentlicht: (2026)
von: Ma, Oubo, et al.
Veröffentlicht: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
The Great Pretender: A Stochasticity Problem in LLM Jailbreak
von: Monteuuis, Jean-Philippe, et al.
Veröffentlicht: (2026)
von: Monteuuis, Jean-Philippe, et al.
Veröffentlicht: (2026)
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
von: Kulkarni, Prashant
Veröffentlicht: (2026)
von: Kulkarni, Prashant
Veröffentlicht: (2026)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
von: Zahran, Noureldin, et al.
Veröffentlicht: (2025)
von: Zahran, Noureldin, et al.
Veröffentlicht: (2025)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
von: Hu, Jiaming, et al.
Veröffentlicht: (2025)
von: Hu, Jiaming, et al.
Veröffentlicht: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
von: Reddy, Pavan, et al.
Veröffentlicht: (2025)
von: Reddy, Pavan, et al.
Veröffentlicht: (2025)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
von: Lin, Zhihao, et al.
Veröffentlicht: (2024)
von: Lin, Zhihao, et al.
Veröffentlicht: (2024)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
von: Hammadia, Taha, et al.
Veröffentlicht: (2026)
von: Hammadia, Taha, et al.
Veröffentlicht: (2026)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks
von: You, Doohee
Veröffentlicht: (2026)
von: You, Doohee
Veröffentlicht: (2026)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
von: Li, Pengcheng, et al.
Veröffentlicht: (2026)
von: Li, Pengcheng, et al.
Veröffentlicht: (2026)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Introducing the Generative Application Firewall (GAF)
von: Farreny, Joan Vendrell, et al.
Veröffentlicht: (2026) -
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025) -
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024) -
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2025) -
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)