RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913059644112896 |
|---|---|
| author | Xiao, Wenjie Tang, Xuehai Zhou, Biyu Hu, Songlin Han, Jizhong |
| author_facet | Xiao, Wenjie Tang, Xuehai Zhou, Biyu Hu, Songlin Han, Jizhong |
| contents | Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_22888 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents Xiao, Wenjie Tang, Xuehai Zhou, Biyu Hu, Songlin Han, Jizhong Cryptography and Security Artificial Intelligence Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering |
| title | RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2604.22888 |