Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Kejia, Zhang, Jiawen, Li, Boheng, Li, Pengcheng, Lou, Jian, Feng, Zunlei, Song, Mingli, Jia, Ruoxi, Zhang, Tianwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025)
by: Ma, Avery, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Mitigating Many-Shot Jailbreaking
by: Ackerman, Christopher M., et al.
Published: (2025)
by: Ackerman, Christopher M., et al.
Published: (2025)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
by: Deng, Gelei, et al.
Published: (2024)
by: Deng, Gelei, et al.
Published: (2024)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
CipherGuard: Compiler-aided Mitigation against Ciphertext Side-channel Attacks
by: Jiang, Ke, et al.
Published: (2025)
by: Jiang, Ke, et al.
Published: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
by: Zhang, Junke, et al.
Published: (2026)
by: Zhang, Junke, et al.
Published: (2026)
Demonstration Attack against In-Context Learning for Code Intelligence
by: Ge, Yifei, et al.
Published: (2024)
by: Ge, Yifei, et al.
Published: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
by: Hu, Jiacong, et al.
Published: (2024)
by: Hu, Jiacong, et al.
Published: (2024)
ObfusBFA: A Holistic Approach to Safeguarding DNNs from Different Types of Bit-Flip Attacks
by: Yan, Xiaobei, et al.
Published: (2025)
by: Yan, Xiaobei, et al.
Published: (2025)
Nearest is Not Dearest: Towards Practical Defense against Quantization-conditioned Backdoor Attacks
by: Li, Boheng, et al.
Published: (2024)
by: Li, Boheng, et al.
Published: (2024)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
Is Difficulty Calibration All We Need? Towards More Practical Membership Inference Attacks
by: He, Yu, et al.
Published: (2024)
by: He, Yu, et al.
Published: (2024)
DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
by: Qiu, Han, et al.
Published: (2020)
by: Qiu, Han, et al.
Published: (2020)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
by: Lin, Runqi, et al.
Published: (2025)
by: Lin, Runqi, et al.
Published: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
by: Hong, Wenjing, et al.
Published: (2026)
by: Hong, Wenjing, et al.
Published: (2026)
Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
by: Chen, Yukun, et al.
Published: (2025)
by: Chen, Yukun, et al.
Published: (2025)
Guardians of the Agentic System: Preventing Many Shots Jailbreak with Agentic System
by: Barua, Saikat, et al.
Published: (2025)
by: Barua, Saikat, et al.
Published: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
by: Teng, Ma, et al.
Published: (2024)
by: Teng, Ma, et al.
Published: (2024)
Mitigating Data Poisoning Attacks to Local Differential Privacy
by: Li, Xiaolin, et al.
Published: (2025)
by: Li, Xiaolin, et al.
Published: (2025)
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
by: Yang, Yulong, et al.
Published: (2024)
by: Yang, Yulong, et al.
Published: (2024)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
Similar Items
-
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
by: Zhang, Jiawen, et al.
Published: (2025) -
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
by: Zhang, Jiawen, et al.
Published: (2025) -
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025) -
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025) -
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)