Lessons from Defending Gemini Against Indirect Prompt Injections
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Chongyang, Lin, Sharon, Song, Shuang, Hayes, Jamie, Shumailov, Ilia, Yona, Itay, Pluto, Juliette, Pappu, Aneesh, Choquette-Choo, Christopher A., Nasr, Milad, Sitawarin, Chawin, Gibson, Gena, Terzis, Andreas, Flynn, John "Four" |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Extracting alignment data in open models
di: Barbero, Federico, et al.
Pubblicazione: (2025)
di: Barbero, Federico, et al.
Pubblicazione: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Buffer Overflow in Mixture of Experts
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Measuring memorization in language models via probabilistic extraction
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Measuring memorization in RLHF for code completion
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
Soft Instruction De-escalation Defense
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
Stealing User Prompts from Mixture of Experts
di: Yona, Itay, et al.
Pubblicazione: (2024)
di: Yona, Itay, et al.
Pubblicazione: (2024)
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
Interpreting the Repeated Token Phenomenon in Large Language Models
di: Yona, Itay, et al.
Pubblicazione: (2025)
di: Yona, Itay, et al.
Pubblicazione: (2025)
Cascading Adversarial Bias from Injection to Distillation in Language Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2025)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2025)
Defeating Prompt Injections by Design
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
Positional Embedding-Aware Activations
di: Shah, Kathan, et al.
Pubblicazione: (2023)
di: Shah, Kathan, et al.
Pubblicazione: (2023)
Operationalizing Contextual Integrity in Privacy-Conscious Assistants
di: Ghalebikesabi, Sahra, et al.
Pubblicazione: (2024)
di: Ghalebikesabi, Sahra, et al.
Pubblicazione: (2024)
PubDef: Defending Against Transfer Attacks From Public Models
di: Sitawarin, Chawin, et al.
Pubblicazione: (2023)
di: Sitawarin, Chawin, et al.
Pubblicazione: (2023)
Auditing Private Prediction
di: Chadha, Karan, et al.
Pubblicazione: (2024)
di: Chadha, Karan, et al.
Pubblicazione: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
di: Steinke, Thomas, et al.
Pubblicazione: (2024)
di: Steinke, Thomas, et al.
Pubblicazione: (2024)
Privacy Auditing of Large Language Models
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
di: Shumailov, Ilia, et al.
Pubblicazione: (2024)
di: Shumailov, Ilia, et al.
Pubblicazione: (2024)
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
di: Lin, Sharon, et al.
Pubblicazione: (2025)
di: Lin, Sharon, et al.
Pubblicazione: (2025)
OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution Shift
di: Li, Lin, et al.
Pubblicazione: (2023)
di: Li, Lin, et al.
Pubblicazione: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
Beyond Slow Signs in High-fidelity Model Extraction
di: Foerster, Hanna, et al.
Pubblicazione: (2024)
di: Foerster, Hanna, et al.
Pubblicazione: (2024)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
Privacy Side Channels in Machine Learning Systems
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
LLMs unlock new paths to monetizing exploits
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
Exploring the limits of strong membership inference attacks on large language models
di: Hayes, Jamie, et al.
Pubblicazione: (2025)
di: Hayes, Jamie, et al.
Pubblicazione: (2025)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
di: Lu, Weikai, et al.
Pubblicazione: (2025)
di: Lu, Weikai, et al.
Pubblicazione: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
di: Gao, Yue, et al.
Pubblicazione: (2023)
di: Gao, Yue, et al.
Pubblicazione: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
di: Wyllie, Sierra, et al.
Pubblicazione: (2024)
di: Wyllie, Sierra, et al.
Pubblicazione: (2024)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Extracting alignment data in open models
di: Barbero, Federico, et al.
Pubblicazione: (2025) -
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025) -
Buffer Overflow in Mixture of Experts
di: Hayes, Jamie, et al.
Pubblicazione: (2024) -
Measuring memorization in language models via probabilistic extraction
di: Hayes, Jamie, et al.
Pubblicazione: (2024) -
Measuring memorization in RLHF for code completion
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)