AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rassul, Yassin H., Rashid, Tarik A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
von: Lee, Yohan, et al.
Veröffentlicht: (2026)
von: Lee, Yohan, et al.
Veröffentlicht: (2026)
Security Attacks on LLM-based Code Completion Tools
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
von: Xu, Naen, et al.
Veröffentlicht: (2026)
von: Xu, Naen, et al.
Veröffentlicht: (2026)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
von: Mohammadi, Bardia, et al.
Veröffentlicht: (2026)
von: Mohammadi, Bardia, et al.
Veröffentlicht: (2026)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
von: Jha, Rishi, et al.
Veröffentlicht: (2026)
von: Jha, Rishi, et al.
Veröffentlicht: (2026)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
von: An, Hengyu, et al.
Veröffentlicht: (2025)
von: An, Hengyu, et al.
Veröffentlicht: (2025)
Multi-use LLM Watermarking and the False Detection Problem
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
OneShield -- the Next Generation of LLM Guardrails
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
von: Wen, Yan, et al.
Veröffentlicht: (2025)
von: Wen, Yan, et al.
Veröffentlicht: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026)
von: Wen, Yule, et al.
Veröffentlicht: (2026)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
von: Wang, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoqing, et al.
Veröffentlicht: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2026)
von: Das, Saswat, et al.
Veröffentlicht: (2026)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2025)
von: Das, Saswat, et al.
Veröffentlicht: (2025)
Searching for Privacy Risks in LLM Agents via Simulation
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
von: Roy, Soham, et al.
Veröffentlicht: (2026)
von: Roy, Soham, et al.
Veröffentlicht: (2026)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Large Language Model Sentinel: LLM Agent for Adversarial Purification
von: Lin, Guang, et al.
Veröffentlicht: (2024)
von: Lin, Guang, et al.
Veröffentlicht: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
Evolving Deception: When Agents Evolve, Deception Wins
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
von: Verma, Rakesh M., et al.
Veröffentlicht: (2024)
von: Verma, Rakesh M., et al.
Veröffentlicht: (2024)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026) -
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026) -
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2024) -
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025) -
Overthinking Loops in Agents: A Structural Risk via MCP Tools
von: Lee, Yohan, et al.
Veröffentlicht: (2026)