Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Yifeng, Wang, Ziming, Deng, Zhaomeng, Yao, Mengyu, Liu, Junlin, Hu, Yutao, Zhang, Ziqi, Guo, Yao, Li, Ding |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Who Moved My Transaction? Uncovering Post-Transaction Auditability Vulnerabilities in Modern Super Apps
di: Liu, Junlin, et al.
Pubblicazione: (2025)
di: Liu, Junlin, et al.
Pubblicazione: (2025)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
di: Lin, Junda, et al.
Pubblicazione: (2026)
di: Lin, Junda, et al.
Pubblicazione: (2026)
I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments When Attackers have Pre-Trained Models
di: Li, Ding, et al.
Pubblicazione: (2024)
di: Li, Ding, et al.
Pubblicazione: (2024)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
di: Lu, Weikai, et al.
Pubblicazione: (2025)
di: Lu, Weikai, et al.
Pubblicazione: (2025)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
di: Weng, Shihao, et al.
Pubblicazione: (2026)
di: Weng, Shihao, et al.
Pubblicazione: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
di: Gong, Guangyu, et al.
Pubblicazione: (2026)
di: Gong, Guangyu, et al.
Pubblicazione: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
Moss: Proxy Model-based Full-Weight Aggregation in Federated Learning with Heterogeneous Models
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
di: Cai, Yifeng, et al.
Pubblicazione: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Connect the Dots: Knowledge Graph-Guided Crawler Attack on Retrieval-Augmented Generation Systems
di: Yao, Mengyu, et al.
Pubblicazione: (2026)
di: Yao, Mengyu, et al.
Pubblicazione: (2026)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion
di: Liu, Ruofan, et al.
Pubblicazione: (2025)
di: Liu, Ruofan, et al.
Pubblicazione: (2025)
Membership Inference Attacks Against Video Large Language Models
di: Song, Wei, et al.
Pubblicazione: (2026)
di: Song, Wei, et al.
Pubblicazione: (2026)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
di: Li, Rongchang, et al.
Pubblicazione: (2024)
di: Li, Rongchang, et al.
Pubblicazione: (2024)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
di: Wang, Yihan, et al.
Pubblicazione: (2025)
di: Wang, Yihan, et al.
Pubblicazione: (2025)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion
di: Tang, Duoxun, et al.
Pubblicazione: (2025)
di: Tang, Duoxun, et al.
Pubblicazione: (2025)
Defending Against Neural Network Model Inversion Attacks via Data Poisoning
di: Zhou, Shuai, et al.
Pubblicazione: (2024)
di: Zhou, Shuai, et al.
Pubblicazione: (2024)
Defending Against Intelligent Attackers at Large Scales
di: Lohn, Andrew J.
Pubblicazione: (2025)
di: Lohn, Andrew J.
Pubblicazione: (2025)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
di: Shah, Harsh
Pubblicazione: (2026)
di: Shah, Harsh
Pubblicazione: (2026)
Secure Distributed Learning for CAVs: Defending Against Gradient Leakage with Leveled Homomorphic Encryption
di: Najjar, Muhammad Ali, et al.
Pubblicazione: (2025)
di: Najjar, Muhammad Ali, et al.
Pubblicazione: (2025)
Defending Against Attack on the Cloned: In-Band Active Man-in-the-Middle Detection for the Signal Protocol
di: Teng, Wil Liam, et al.
Pubblicazione: (2024)
di: Teng, Wil Liam, et al.
Pubblicazione: (2024)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
di: Xu, Wenpeng
Pubblicazione: (2026)
di: Xu, Wenpeng
Pubblicazione: (2026)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
di: He, Yu, et al.
Pubblicazione: (2026)
di: He, Yu, et al.
Pubblicazione: (2026)
Injection Attacks Against End-to-End Encrypted Applications
di: Fábrega, Andrés, et al.
Pubblicazione: (2024)
di: Fábrega, Andrés, et al.
Pubblicazione: (2024)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Who Moved My Transaction? Uncovering Post-Transaction Auditability Vulnerabilities in Modern Super Apps
di: Liu, Junlin, et al.
Pubblicazione: (2025) -
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
di: Lin, Junda, et al.
Pubblicazione: (2026) -
I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
di: Cai, Yifeng, et al.
Pubblicazione: (2025) -
TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments When Attackers have Pre-Trained Models
di: Li, Ding, et al.
Pubblicazione: (2024) -
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)