Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Hanjiang, Robey, Alexander, Liu, Changliu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Jailbreaking Attack against Multimodal Large Language Model
by: Niu, Zhenxing, et al.
Published: (2024)
by: Niu, Zhenxing, et al.
Published: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
by: Li, Xiao, et al.
Published: (2024)
by: Li, Xiao, et al.
Published: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
by: Li, Qizhang, et al.
Published: (2024)
by: Li, Qizhang, et al.
Published: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
by: Rastogi, Saksham, et al.
Published: (2024)
by: Rastogi, Saksham, et al.
Published: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
by: Fang, Zhicheng, et al.
Published: (2026)
by: Fang, Zhicheng, et al.
Published: (2026)
Jailbreaking with Universal Multi-Prompts
by: Hsu, Yu-Ling, et al.
Published: (2025)
by: Hsu, Yu-Ling, et al.
Published: (2025)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
by: Fu, Wenjie, et al.
Published: (2023)
by: Fu, Wenjie, et al.
Published: (2023)
A StrongREJECT for Empty Jailbreaks
by: Souly, Alexandra, et al.
Published: (2024)
by: Souly, Alexandra, et al.
Published: (2024)
Adaptive Probe-based Steering for Robust LLM Jailbreaking
by: Chen, Junxi, et al.
Published: (2026)
by: Chen, Junxi, et al.
Published: (2026)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
by: Nguyen, Khanh, et al.
Published: (2025)
by: Nguyen, Khanh, et al.
Published: (2025)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
by: Jia, Xiaojun, et al.
Published: (2024)
by: Jia, Xiaojun, et al.
Published: (2024)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
by: Dotsinski, Asen, et al.
Published: (2026)
by: Dotsinski, Asen, et al.
Published: (2026)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
by: Kim, Taeyoun, et al.
Published: (2024)
by: Kim, Taeyoun, et al.
Published: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Similar Items
-
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024) -
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025) -
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024) -
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
by: Jiang, Yifan, et al.
Published: (2024) -
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)