The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Fuente:
arXiv
Saved in:
| Main Authors: | Nasr, Milad, Carlini, Nicholas, Sitawarin, Chawin, Schulhoff, Sander V., Hayes, Jamie, Ilie, Michael, Pluto, Juliette, Song, Shuang, Chaudhari, Harsh, Shumailov, Ilia, Thakurta, Abhradeep, Xiao, Kai Yuanqing, Terzis, Andreas, Tramèr, Florian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025)
by: Debenedetti, Edoardo, et al.
Published: (2025)
Hush! Protecting Secrets During Model Training: An Indistinguishability Approach
by: Ganesh, Arun, et al.
Published: (2025)
by: Ganesh, Arun, et al.
Published: (2025)
Positional Embedding-Aware Activations
by: Shah, Kathan, et al.
Published: (2023)
by: Shah, Kathan, et al.
Published: (2023)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Query-Based Adversarial Prompt Generation
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Extracting alignment data in open models
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Reinforced Generation of Combinatorial Structures: Ramsey Numbers
by: Nagda, Ansh, et al.
Published: (2026)
by: Nagda, Ansh, et al.
Published: (2026)
Reinforced Generation of Combinatorial Structures: Hardness of Approximation
by: Nagda, Ansh, et al.
Published: (2025)
by: Nagda, Ansh, et al.
Published: (2025)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution Shift
by: Li, Lin, et al.
Published: (2023)
by: Li, Lin, et al.
Published: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
StruQ: Defending Against Prompt Injection with Structured Queries
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
On Design Principles for Private Adaptive Optimizers
by: Ganesh, Arun, et al.
Published: (2025)
by: Ganesh, Arun, et al.
Published: (2025)
InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
by: Vinod, Vishnu, et al.
Published: (2025)
by: Vinod, Vishnu, et al.
Published: (2025)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
Optimal Rates for $O(1)$-Smooth DP-SCO with a Single Epoch and Large Batches
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Measuring memorization in RLHF for code completion
by: Pappu, Aneesh, et al.
Published: (2024)
by: Pappu, Aneesh, et al.
Published: (2024)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
by: Zhao, Yunhan, et al.
Published: (2025)
by: Zhao, Yunhan, et al.
Published: (2025)
PubDef: Defending Against Transfer Attacks From Public Models
by: Sitawarin, Chawin, et al.
Published: (2023)
by: Sitawarin, Chawin, et al.
Published: (2023)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Differentially private survey research
by: Georgina Evans, et al.
Published: (2024)
by: Georgina Evans, et al.
Published: (2024)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Privacy Amplification for Matrix Mechanisms
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Differentially Private Parameter-Efficient Fine-tuning for Large ASR Models
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
Training Large ASR Encoders with Differential Privacy
by: Chauhan, Geeticka, et al.
Published: (2024)
by: Chauhan, Geeticka, et al.
Published: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
by: Rando, Javier, et al.
Published: (2023)
by: Rando, Javier, et al.
Published: (2023)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Similar Items
-
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025) -
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025) -
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025) -
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026) -
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025)