Soft Instruction De-escalation Defense
Fuente:
arXiv
Saved in:
| Main Authors: | Walter, Nils Philipp, Sitawarin, Chawin, Hayes, Jamie, Stutz, David, Shumailov, Ilia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
by: Gao, Yue, et al.
Published: (2023)
by: Gao, Yue, et al.
Published: (2023)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
ceLLMate: Sandboxing Browser AI Agents
by: Meng, Luoxi, et al.
Published: (2025)
by: Meng, Luoxi, et al.
Published: (2025)
Locking Machine Learning Models into Hardware
by: Clifford, Eleanor, et al.
Published: (2024)
by: Clifford, Eleanor, et al.
Published: (2024)
PubDef: Defending Against Transfer Attacks From Public Models
by: Sitawarin, Chawin, et al.
Published: (2023)
by: Sitawarin, Chawin, et al.
Published: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
Watermarking Needs Input Repetition Masking
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
by: Rinberg, Roy, et al.
Published: (2025)
by: Rinberg, Roy, et al.
Published: (2025)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
by: Clifford, Eleanor, et al.
Published: (2022)
by: Clifford, Eleanor, et al.
Published: (2022)
When Vision Fails: Text Attacks Against ViT and OCR
by: Boucher, Nicholas, et al.
Published: (2023)
by: Boucher, Nicholas, et al.
Published: (2023)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
StruQ: Defending Against Prompt Injection with Structured Queries
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Machine Learning Models Have a Supply Chain Problem
by: Meiklejohn, Sarah, et al.
Published: (2025)
by: Meiklejohn, Sarah, et al.
Published: (2025)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)
by: Thudi, Anvith, et al.
Published: (2023)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography
by: Shumailov, Ilia, et al.
Published: (2025)
by: Shumailov, Ilia, et al.
Published: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
Architectural Neural Backdoors from First Principles
by: Langford, Harry, et al.
Published: (2024)
by: Langford, Harry, et al.
Published: (2024)
Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks
by: Fu, Yanzhang, et al.
Published: (2026)
by: Fu, Yanzhang, et al.
Published: (2026)
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
by: elShehaby, Mohamed, et al.
Published: (2026)
by: elShehaby, Mohamed, et al.
Published: (2026)
Exploring the limits of strong membership inference attacks on large language models
by: Hayes, Jamie, et al.
Published: (2025)
by: Hayes, Jamie, et al.
Published: (2025)
SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
by: Luo, Chenxiang, et al.
Published: (2025)
by: Luo, Chenxiang, et al.
Published: (2025)
Similar Items
-
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024) -
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025) -
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024) -
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024) -
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)