Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Sandoval, Gustavo, Fenchenko, Denys, Chen, Junyao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
por: Yang, Xiaoxue, et al.
Publicado: (2025)
por: Yang, Xiaoxue, et al.
Publicado: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
por: Hossain, S M Asif, et al.
Publicado: (2025)
por: Hossain, S M Asif, et al.
Publicado: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
por: Yin, Chenlong, et al.
Publicado: (2026)
por: Yin, Chenlong, et al.
Publicado: (2026)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
por: Debenedetti, Edoardo, et al.
Publicado: (2024)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
por: Kuo, Kevin, et al.
Publicado: (2026)
por: Kuo, Kevin, et al.
Publicado: (2026)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
por: Zhang, Shuhao, et al.
Publicado: (2026)
por: Zhang, Shuhao, et al.
Publicado: (2026)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023)
por: Yan, Jun, et al.
Publicado: (2023)
Fine-Tuning Personalization in Federated Learning to Mitigate Adversarial Clients
por: Allouah, Youssef, et al.
Publicado: (2024)
por: Allouah, Youssef, et al.
Publicado: (2024)
An Early Categorization of Prompt Injection Attacks on Large Language Models
por: Rossi, Sippo, et al.
Publicado: (2024)
por: Rossi, Sippo, et al.
Publicado: (2024)
Lorica: A Synergistic Fine-Tuning Framework for Advancing Personalized Adversarial Robustness
por: Qi, Tianyu, et al.
Publicado: (2025)
por: Qi, Tianyu, et al.
Publicado: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023)
por: Liu, Yupei, et al.
Publicado: (2023)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
por: Panebianco, Francesco, et al.
Publicado: (2025)
por: Panebianco, Francesco, et al.
Publicado: (2025)
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions
por: Li, Wenjuan, et al.
Publicado: (2026)
por: Li, Wenjuan, et al.
Publicado: (2026)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
por: Thakkar, Janvi, et al.
Publicado: (2023)
por: Thakkar, Janvi, et al.
Publicado: (2023)
Cascading Adversarial Bias from Injection to Distillation in Language Models
por: Chaudhari, Harsh, et al.
Publicado: (2025)
por: Chaudhari, Harsh, et al.
Publicado: (2025)
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
por: Pandey, Rohan, et al.
Publicado: (2026)
por: Pandey, Rohan, et al.
Publicado: (2026)
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
por: Xin, Yuan, et al.
Publicado: (2023)
por: Xin, Yuan, et al.
Publicado: (2023)
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
por: elShehaby, Mohamed, et al.
Publicado: (2026)
por: elShehaby, Mohamed, et al.
Publicado: (2026)
IDEA: Invariant Defense for Graph Adversarial Robustness
por: Tao, Shuchang, et al.
Publicado: (2023)
por: Tao, Shuchang, et al.
Publicado: (2023)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
por: Maiorano, Alexandre Cristovão
Publicado: (2026)
por: Maiorano, Alexandre Cristovão
Publicado: (2026)
SecAlign: Defending Against Prompt Injection with Preference Optimization
por: Chen, Sizhe, et al.
Publicado: (2024)
por: Chen, Sizhe, et al.
Publicado: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
por: Khachaturov, David, et al.
Publicado: (2025)
por: Khachaturov, David, et al.
Publicado: (2025)
Enhancing the "Immunity" of Mixture-of-Experts Networks for Adversarial Defense
por: Han, Qiao, et al.
Publicado: (2024)
por: Han, Qiao, et al.
Publicado: (2024)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
por: Yeh, Cheng-Han, et al.
Publicado: (2024)
por: Yeh, Cheng-Han, et al.
Publicado: (2024)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
por: Zou, Wei, et al.
Publicado: (2025)
por: Zou, Wei, et al.
Publicado: (2025)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
por: Lan, Wenhao, et al.
Publicado: (2026)
por: Lan, Wenhao, et al.
Publicado: (2026)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
por: Liu, Mingrui, et al.
Publicado: (2026)
por: Liu, Mingrui, et al.
Publicado: (2026)
Preventing Prompt Injection with Type-Directed Privilege Separation
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
por: Konrad, Phongsakon Mark, et al.
Publicado: (2026)
por: Konrad, Phongsakon Mark, et al.
Publicado: (2026)
Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses
por: Wang, Yongyu
Publicado: (2025)
por: Wang, Yongyu
Publicado: (2025)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
por: Li, Rongchang, et al.
Publicado: (2024)
por: Li, Rongchang, et al.
Publicado: (2024)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
por: Clop, Cody, et al.
Publicado: (2024)
por: Clop, Cody, et al.
Publicado: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
por: Wu, Yuanwei, et al.
Publicado: (2023)
por: Wu, Yuanwei, et al.
Publicado: (2023)
Efficient Differentially Private Fine-Tuning of Diffusion Models
por: Liu, Jing, et al.
Publicado: (2024)
por: Liu, Jing, et al.
Publicado: (2024)
Design Patterns for Securing LLM Agents against Prompt Injections
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
por: Beurer-Kellner, Luca, et al.
Publicado: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
por: Shi, Chongyang, et al.
Publicado: (2025)
por: Shi, Chongyang, et al.
Publicado: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025)
por: Zizzo, Giulio, et al.
Publicado: (2025)
Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
por: Nguyen, Quan Minh, et al.
Publicado: (2026)
Ejemplares similares
-
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
por: Yang, Xiaoxue, et al.
Publicado: (2025) -
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
por: Hossain, S M Asif, et al.
Publicado: (2025) -
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
por: Yin, Chenlong, et al.
Publicado: (2026) -
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
por: Debenedetti, Edoardo, et al.
Publicado: (2024) -
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)