Defending Jailbreak Prompts via In-Context Adversarial Game
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yujun, Han, Yufei, Zhuang, Haomin, Guo, Kehan, Liang, Zhenwen, Bao, Hongyan, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AIRGuard: Guarding Agent Actions with Runtime Authority Control
von: Qin, Suliu, et al.
Veröffentlicht: (2026)
von: Qin, Suliu, et al.
Veröffentlicht: (2026)
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
von: Fu, Shaopeng, et al.
Veröffentlicht: (2025)
von: Fu, Shaopeng, et al.
Veröffentlicht: (2025)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
VLMGuard: Defending VLMs against Malicious Prompts via Unlabeled Data
von: Du, Xuefeng, et al.
Veröffentlicht: (2024)
von: Du, Xuefeng, et al.
Veröffentlicht: (2024)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
Lurking in the shadows: Unveiling Stealthy Backdoor Attacks against Personalized Federated Learning
von: Lyu, Xiaoting, et al.
Veröffentlicht: (2024)
von: Lyu, Xiaoting, et al.
Veröffentlicht: (2024)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
MEA-Defender: A Robust Watermark against Model Extraction Attack
von: Lv, Peizhuo, et al.
Veröffentlicht: (2024)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
von: Zheng, Meixi, et al.
Veröffentlicht: (2025)
von: Zheng, Meixi, et al.
Veröffentlicht: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
von: Li, Xuan, et al.
Veröffentlicht: (2023)
von: Li, Xuan, et al.
Veröffentlicht: (2023)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
von: Jin, Zhibo, et al.
Veröffentlicht: (2024)
von: Jin, Zhibo, et al.
Veröffentlicht: (2024)
Stealix: Model Stealing via Prompt Evolution
von: Zhuang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhuang, Zhixiong, et al.
Veröffentlicht: (2025)
Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
RADAR: Defending RAG Dynamically against Retrieval Corruption
von: Chen, Ziyuan, et al.
Veröffentlicht: (2026)
von: Chen, Ziyuan, et al.
Veröffentlicht: (2026)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
von: Yue, Murong, et al.
Veröffentlicht: (2025)
von: Yue, Murong, et al.
Veröffentlicht: (2025)
Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise
von: Lian, Puwei, et al.
Veröffentlicht: (2026)
von: Lian, Puwei, et al.
Veröffentlicht: (2026)
Enhancing Membership Inference Attacks on Diffusion Models from a Frequency-Domain Perspective
von: Lian, Puwei, et al.
Veröffentlicht: (2025)
von: Lian, Puwei, et al.
Veröffentlicht: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
von: Zizzo, Giulio, et al.
Veröffentlicht: (2025)
von: Zizzo, Giulio, et al.
Veröffentlicht: (2025)
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
Soften to Defend: Towards Adversarial Robustness via Self-Guided Label Refinement
von: Yu, Daiwei, et al.
Veröffentlicht: (2024)
von: Yu, Daiwei, et al.
Veröffentlicht: (2024)
Game-Theoretic Defenses for Robust Conformal Prediction Against Adversarial Attacks in Medical Imaging
von: Luo, Rui, et al.
Veröffentlicht: (2024)
von: Luo, Rui, et al.
Veröffentlicht: (2024)
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
Revisiting Label Inference Attacks in Vertical Federated Learning: Why They Are Vulnerable and How to Defend
von: Liu, Yige, et al.
Veröffentlicht: (2026)
von: Liu, Yige, et al.
Veröffentlicht: (2026)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
von: Xie, Zhixin, et al.
Veröffentlicht: (2026)
von: Xie, Zhixin, et al.
Veröffentlicht: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AIRGuard: Guarding Agent Actions with Runtime Authority Control
von: Qin, Suliu, et al.
Veröffentlicht: (2026) -
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
von: Fu, Shaopeng, et al.
Veröffentlicht: (2025) -
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024) -
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023) -
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)