ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Xingwei, Lin, Wenhao, Cao, Sicong, Yu, Jiahao, Huang, Renke, Xue, Lei, Wu, Chunming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
von: Cui, Tiehan, et al.
Veröffentlicht: (2025)
von: Cui, Tiehan, et al.
Veröffentlicht: (2025)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
The Echo Chamber Multi-Turn LLM Jailbreak
von: Alobaid, Ahmad, et al.
Veröffentlicht: (2026)
von: Alobaid, Ahmad, et al.
Veröffentlicht: (2026)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2025)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
MAS-SZZ: Multi-Agentic SZZ Algorithm for Vulnerability-Inducing Commit Identification
von: Cao, Sicong, et al.
Veröffentlicht: (2026)
von: Cao, Sicong, et al.
Veröffentlicht: (2026)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI Ecosystem
von: Gao, Xingan, et al.
Veröffentlicht: (2025)
von: Gao, Xingan, et al.
Veröffentlicht: (2025)
PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning
von: Li, Shenghui, et al.
Veröffentlicht: (2024)
von: Li, Shenghui, et al.
Veröffentlicht: (2024)
Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks
von: You, Doohee
Veröffentlicht: (2026)
von: You, Doohee
Veröffentlicht: (2026)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
SAB:A Stealing and Robust Backdoor Attack based on Steganographic Algorithm against Federated Learning
von: Xu, Weida, et al.
Veröffentlicht: (2024)
von: Xu, Weida, et al.
Veröffentlicht: (2024)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
von: Lu, Lin, et al.
Veröffentlicht: (2024)
von: Lu, Lin, et al.
Veröffentlicht: (2024)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025) -
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
von: Cui, Tiehan, et al.
Veröffentlicht: (2025) -
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024) -
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026) -
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)