RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Wenjie, Tang, Xuehai, Zhou, Biyu, Hu, Songlin, Han, Jizhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
by: Lv, Lijia, et al.
Published: (2026)
by: Lv, Lijia, et al.
Published: (2026)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026)
by: Liu, Shi, et al.
Published: (2026)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
by: Lv, Lijia, et al.
Published: (2024)
by: Lv, Lijia, et al.
Published: (2024)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
by: Langiu, Alessio
Published: (2026)
by: Langiu, Alessio
Published: (2026)
AIRGuard: Guarding Agent Actions with Runtime Authority Control
by: Qin, Suliu, et al.
Published: (2026)
by: Qin, Suliu, et al.
Published: (2026)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection
by: Xue, Yinuo, et al.
Published: (2025)
by: Xue, Yinuo, et al.
Published: (2025)
LoopTrap: Termination Poisoning Attacks on LLM Agents
by: Xu, Huiyu, et al.
Published: (2026)
by: Xu, Huiyu, et al.
Published: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
by: Wei, Qianshan, et al.
Published: (2025)
by: Wei, Qianshan, et al.
Published: (2025)
X-Guard: Multilingual Guard Agent for Content Moderation
by: Upadhayay, Bibek, et al.
Published: (2025)
by: Upadhayay, Bibek, et al.
Published: (2025)
SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents
by: Ouyang, Yipeng, et al.
Published: (2026)
by: Ouyang, Yipeng, et al.
Published: (2026)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
Human-Imperceptible Retrieval Poisoning Attacks in LLM-Powered Applications
by: Zhang, Quan, et al.
Published: (2024)
by: Zhang, Quan, et al.
Published: (2024)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
by: Hou, Yinghan, et al.
Published: (2026)
by: Hou, Yinghan, et al.
Published: (2026)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
by: Zhao, Wei, et al.
Published: (2026)
by: Zhao, Wei, et al.
Published: (2026)
Poison to Detect: Detection of Targeted Overfitting in Federated Learning
by: Mestari, Soumia Zohra El, et al.
Published: (2025)
by: Mestari, Soumia Zohra El, et al.
Published: (2025)
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
by: Wang, Kaixiang, et al.
Published: (2026)
by: Wang, Kaixiang, et al.
Published: (2026)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
by: Zhang, Jinchuan, et al.
Published: (2025)
by: Zhang, Jinchuan, et al.
Published: (2025)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
by: Yang, Xikang, et al.
Published: (2025)
by: Yang, Xikang, et al.
Published: (2025)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
by: Chen, Zhihao, et al.
Published: (2026)
by: Chen, Zhihao, et al.
Published: (2026)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
by: Zhou, Zhaojiacheng
Published: (2026)
by: Zhou, Zhaojiacheng
Published: (2026)
TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents
by: Liu, Yibing, et al.
Published: (2026)
by: Liu, Yibing, et al.
Published: (2026)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
by: Chen, Jizhou, et al.
Published: (2025)
by: Chen, Jizhou, et al.
Published: (2025)
SkillTester: Benchmarking Utility and Security of Agent Skills
by: Wang, Leye, et al.
Published: (2026)
by: Wang, Leye, et al.
Published: (2026)
Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
by: Kereopa-Yorke, Ben, et al.
Published: (2026)
by: Kereopa-Yorke, Ben, et al.
Published: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
by: Lee, Taegyeong, et al.
Published: (2025)
by: Lee, Taegyeong, et al.
Published: (2025)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
by: Qiao, Yuxuan, et al.
Published: (2025)
by: Qiao, Yuxuan, et al.
Published: (2025)
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
by: Jia, Xiaojun, et al.
Published: (2026)
by: Jia, Xiaojun, et al.
Published: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Toward Polymorphic Backdoor against Semantic Communication via Intensity-Based Poisoning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
CoT-Guard: Small Models for Strong Monitoring
by: Diwan, Nirav, et al.
Published: (2026)
by: Diwan, Nirav, et al.
Published: (2026)
Poisoned Acoustics
by: Dahme, Harrison
Published: (2026)
by: Dahme, Harrison
Published: (2026)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
by: Hasan, Md. Mehedi, et al.
Published: (2025)
by: Hasan, Md. Mehedi, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Similar Items
-
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
by: Lv, Lijia, et al.
Published: (2026) -
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026) -
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
by: Lv, Lijia, et al.
Published: (2024) -
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
by: Yang, Xikang, et al.
Published: (2024) -
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)