RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Wenjie, Tang, Xuehai, Zhou, Biyu, Hu, Songlin, Han, Jizhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913059644112896
author Xiao, Wenjie
Tang, Xuehai
Zhou, Biyu
Hu, Songlin
Han, Jizhong
author_facet Xiao, Wenjie
Tang, Xuehai
Zhou, Biyu
Hu, Songlin
Han, Jizhong
contents Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering
format Preprint
id arxiv_https___arxiv_org_abs_2604_22888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
Xiao, Wenjie
Tang, Xuehai
Zhou, Biyu
Hu, Songlin
Han, Jizhong
Cryptography and Security
Artificial Intelligence
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering
title RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2604.22888