RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhong, Peter Yong, Chen, Siyuan, Wang, Ruiqi, McCall, McKenna, Titzer, Ben L., Miller, Heather, Gibbons, Phillip B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
di: Weng, Shihao, et al.
Pubblicazione: (2026)
di: Weng, Shihao, et al.
Pubblicazione: (2026)
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
di: Lin, Junda, et al.
Pubblicazione: (2026)
di: Lin, Junda, et al.
Pubblicazione: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Scaling up the Banded Matrix Factorization Mechanism for Differentially Private ML
di: McKenna, Ryan
Pubblicazione: (2024)
di: McKenna, Ryan
Pubblicazione: (2024)
InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split Inference
di: Deng, Ruijun, et al.
Pubblicazione: (2025)
di: Deng, Ruijun, et al.
Pubblicazione: (2025)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
di: Wang, Yihan, et al.
Pubblicazione: (2025)
di: Wang, Yihan, et al.
Pubblicazione: (2025)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
di: Lu, Weikai, et al.
Pubblicazione: (2025)
di: Lu, Weikai, et al.
Pubblicazione: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents
di: Crawford, Brian, et al.
Pubblicazione: (2026)
di: Crawford, Brian, et al.
Pubblicazione: (2026)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
di: Shah, Harsh
Pubblicazione: (2026)
di: Shah, Harsh
Pubblicazione: (2026)
Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
di: Li, Rongchang, et al.
Pubblicazione: (2024)
di: Li, Rongchang, et al.
Pubblicazione: (2024)
Secure Distributed Learning for CAVs: Defending Against Gradient Leakage with Leveled Homomorphic Encryption
di: Najjar, Muhammad Ali, et al.
Pubblicazione: (2025)
di: Najjar, Muhammad Ali, et al.
Pubblicazione: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
di: Gong, Guangyu, et al.
Pubblicazione: (2026)
di: Gong, Guangyu, et al.
Pubblicazione: (2026)
DeepLeak: Privacy Enhancing Hardening of Model Explanations Against Membership Leakage
di: Hmida, Firas Ben, et al.
Pubblicazione: (2026)
di: Hmida, Firas Ben, et al.
Pubblicazione: (2026)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
di: Xin, Yuan, et al.
Pubblicazione: (2026)
di: Xin, Yuan, et al.
Pubblicazione: (2026)
Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines
di: Huang, Tao, et al.
Pubblicazione: (2026)
di: Huang, Tao, et al.
Pubblicazione: (2026)
Exploiting Leakage in Password Managers via Injection Attacks
di: Fábrega, Andrés, et al.
Pubblicazione: (2024)
di: Fábrega, Andrés, et al.
Pubblicazione: (2024)
Releasing Large-Scale Human Mobility Histograms with Differential Privacy
di: Bian, Christopher, et al.
Pubblicazione: (2024)
di: Bian, Christopher, et al.
Pubblicazione: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
CRFU: Compressive Representation Forgetting Against Privacy Leakage on Machine Unlearning
di: Wang, Weiqi, et al.
Pubblicazione: (2025)
di: Wang, Weiqi, et al.
Pubblicazione: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025)
di: Wang, Reachal, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026) -
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
di: Weng, Shihao, et al.
Pubblicazione: (2026) -
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025) -
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025) -
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)