Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Shuhao, Li, Jiarui, Cao, Qi, Zhang, Ruiyi, Xie, Pengtao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
di: Xue, Eric, et al.
Pubblicazione: (2025)
di: Xue, Eric, et al.
Pubblicazione: (2025)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
di: Cao, Qi, et al.
Pubblicazione: (2026)
di: Cao, Qi, et al.
Pubblicazione: (2026)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
di: Liu, Mingrui, et al.
Pubblicazione: (2026)
di: Liu, Mingrui, et al.
Pubblicazione: (2026)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
di: Sandoval, Gustavo, et al.
Pubblicazione: (2025)
di: Sandoval, Gustavo, et al.
Pubblicazione: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
di: Maiorano, Alexandre Cristovão
Pubblicazione: (2026)
di: Maiorano, Alexandre Cristovão
Pubblicazione: (2026)
FilterFL: Knowledge Filtering-based Data-Free Backdoor Defense for Federated Learning
di: Yang, Yanxin, et al.
Pubblicazione: (2023)
di: Yang, Yanxin, et al.
Pubblicazione: (2023)
Preventing Prompt Injection with Type-Directed Privilege Separation
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
IDEA: Invariant Defense for Graph Adversarial Robustness
di: Tao, Shuchang, et al.
Pubblicazione: (2023)
di: Tao, Shuchang, et al.
Pubblicazione: (2023)
Learning to Look Benign: Targeted Evasion of Malware Detectors via API Import Injection
di: Dautartas, Juozas, et al.
Pubblicazione: (2026)
di: Dautartas, Juozas, et al.
Pubblicazione: (2026)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Design Patterns for Securing LLM Agents against Prompt Injections
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks
di: Fu, Yanzhang, et al.
Pubblicazione: (2026)
di: Fu, Yanzhang, et al.
Pubblicazione: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
Certified Defense on the Fairness of Graph Neural Networks
di: Dong, Yushun, et al.
Pubblicazione: (2023)
di: Dong, Yushun, et al.
Pubblicazione: (2023)
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
di: Pasquini, Dario, et al.
Pubblicazione: (2024)
di: Pasquini, Dario, et al.
Pubblicazione: (2024)
Adaptive Deception Framework with Behavioral Analysis for Enhanced Cybersecurity Defense
di: AL-Zahrani, Basil Abdullah
Pubblicazione: (2025)
di: AL-Zahrani, Basil Abdullah
Pubblicazione: (2025)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
di: Li, Rongchang, et al.
Pubblicazione: (2024)
di: Li, Rongchang, et al.
Pubblicazione: (2024)
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
di: Sahabandu, Dinuka, et al.
Pubblicazione: (2024)
di: Sahabandu, Dinuka, et al.
Pubblicazione: (2024)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
di: Clop, Cody, et al.
Pubblicazione: (2024)
di: Clop, Cody, et al.
Pubblicazione: (2024)
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
Combining Machine Learning Defenses without Conflicts
di: Duddu, Vasisht, et al.
Pubblicazione: (2024)
di: Duddu, Vasisht, et al.
Pubblicazione: (2024)
Evaluations of Machine Learning Privacy Defenses are Misleading
di: Aerni, Michael, et al.
Pubblicazione: (2024)
di: Aerni, Michael, et al.
Pubblicazione: (2024)
BeniFul: Backdoor Defense via Middle Feature Analysis for Deep Neural Networks
di: Li, Xinfu, et al.
Pubblicazione: (2024)
di: Li, Xinfu, et al.
Pubblicazione: (2024)
ARBoids: Adaptive Residual Reinforcement Learning With Boids Model for Cooperative Multi-USV Target Defense
di: Tao, Jiyue, et al.
Pubblicazione: (2025)
di: Tao, Jiyue, et al.
Pubblicazione: (2025)
Data Reconstruction Attacks and Defenses: A Systematic Evaluation
di: Liu, Sheng, et al.
Pubblicazione: (2024)
di: Liu, Sheng, et al.
Pubblicazione: (2024)
Robustness Inspired Graph Backdoor Defense
di: Zhang, Zhiwei, et al.
Pubblicazione: (2024)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2024)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Kaiyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
di: Xue, Eric, et al.
Pubblicazione: (2025) -
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
di: Zhan, Qiusi, et al.
Pubblicazione: (2025) -
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025) -
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
di: Cao, Qi, et al.
Pubblicazione: (2026) -
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026)