Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Lecheng, Li, Ruizhe, Han, Xicheng, Li, Wenxi, Wang, Binwu, Wang, Longyue, Lyu, Chenyang, Chen, Guanhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
by: Rassul, Yassin H., et al.
Published: (2026)
by: Rassul, Yassin H., et al.
Published: (2026)
Security Attacks on LLM-based Code Completion Tools
by: Cheng, Wen, et al.
Published: (2024)
by: Cheng, Wen, et al.
Published: (2024)
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026)
by: Hu, Yuepeng, et al.
Published: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
by: Du, Chenghao, et al.
Published: (2025)
by: Du, Chenghao, et al.
Published: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
Imprompter: Tricking LLM Agents into Improper Tool Use
by: Fu, Xiaohan, et al.
Published: (2024)
by: Fu, Xiaohan, et al.
Published: (2024)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents
by: Li, Zichuan, et al.
Published: (2025)
by: Li, Zichuan, et al.
Published: (2025)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
by: Mohammadi, Bardia, et al.
Published: (2026)
by: Mohammadi, Bardia, et al.
Published: (2026)
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
by: Sandoval, Aaron, et al.
Published: (2025)
by: Sandoval, Aaron, et al.
Published: (2025)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
by: Lee, Yohan, et al.
Published: (2026)
by: Lee, Yohan, et al.
Published: (2026)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
by: Lin, Junda, et al.
Published: (2026)
by: Lin, Junda, et al.
Published: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
by: Zhan, Qiusi, et al.
Published: (2024)
by: Zhan, Qiusi, et al.
Published: (2024)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
by: An, Hengyu, et al.
Published: (2025)
by: An, Hengyu, et al.
Published: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents
by: Chinaei, Mohammad Hossein
Published: (2026)
by: Chinaei, Mohammad Hossein
Published: (2026)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
by: Li, Jing-Jing, et al.
Published: (2025)
by: Li, Jing-Jing, et al.
Published: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation
by: Ye, Hengkai, et al.
Published: (2026)
by: Ye, Hengkai, et al.
Published: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
by: Qiao, Yuxuan, et al.
Published: (2025)
by: Qiao, Yuxuan, et al.
Published: (2025)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
by: Tshimula, Jean Marie, et al.
Published: (2024)
by: Tshimula, Jean Marie, et al.
Published: (2024)
A Broad Comparative Evaluation of Software Debloating Tools
by: Brown, Michael D., et al.
Published: (2023)
by: Brown, Michael D., et al.
Published: (2023)
Prompt Injection Attack to Tool Selection in LLM Agents
by: Shi, Jiawen, et al.
Published: (2025)
by: Shi, Jiawen, et al.
Published: (2025)
Triad: Trusted Timestamps in Untrusted Environments
by: Fernandez, Gabriel P., et al.
Published: (2023)
by: Fernandez, Gabriel P., et al.
Published: (2023)
From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
Attestable Builds: Compiling Verifiable Binaries on Untrusted Systems using Trusted Execution Environments
by: Hugenroth, Daniel, et al.
Published: (2025)
by: Hugenroth, Daniel, et al.
Published: (2025)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)
by: Liang, Jiacheng, et al.
Published: (2024)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
by: Zhou, Zhenhong, et al.
Published: (2026)
by: Zhou, Zhenhong, et al.
Published: (2026)
Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents
by: Li, Zongwei, et al.
Published: (2026)
by: Li, Zongwei, et al.
Published: (2026)
OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents
by: Li, Frank
Published: (2026)
by: Li, Frank
Published: (2026)
Similar Items
-
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
by: Rassul, Yassin H., et al.
Published: (2026) -
Security Attacks on LLM-based Code Completion Tools
by: Cheng, Wen, et al.
Published: (2024) -
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026) -
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026) -
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
by: Du, Chenghao, et al.
Published: (2025)