Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schmotz, David, Beurer-Kellner, Luca, Abdelnabi, Sahar, Andriushchenko, Maksym
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914350603698176
author Schmotz, David
Beurer-Kellner, Luca
Abdelnabi, Sahar
Andriushchenko, Maksym
author_facet Schmotz, David
Beurer-Kellner, Luca
Abdelnabi, Sahar
Andriushchenko, Maksym
contents LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized third-party code, knowledge, and instructions. Although this can extend agent capabilities to new domains, it creates an increasingly complex agent supply chain, offering new surfaces for prompt injection attacks. We identify skill-based prompt injection as a significant threat and introduce SkillInject, a benchmark evaluating the susceptibility of widely-used LLM agents to injections through skill files. SkillInject contains 202 injection-task pairs with attacks ranging from obviously malicious injections to subtle, context-dependent attacks hidden in otherwise legitimate instructions. We evaluate frontier LLMs on SkillInject, measuring both security in terms of harmful instruction avoidance and utility in terms of legitimate instruction compliance. Our results show that today's agents are highly vulnerable with up to 80% attack success rate with frontier models, often executing extremely harmful instructions including data exfiltration, destructive action, and ransomware-like behavior. They furthermore suggest that this problem will not be solved through model scaling or simple input filtering, but that robust agent security will require context-aware authorization frameworks. Our benchmark is available at https://www.skill-inject.com/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20156
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
Schmotz, David
Beurer-Kellner, Luca
Abdelnabi, Sahar
Andriushchenko, Maksym
Cryptography and Security
Machine Learning
LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized third-party code, knowledge, and instructions. Although this can extend agent capabilities to new domains, it creates an increasingly complex agent supply chain, offering new surfaces for prompt injection attacks. We identify skill-based prompt injection as a significant threat and introduce SkillInject, a benchmark evaluating the susceptibility of widely-used LLM agents to injections through skill files. SkillInject contains 202 injection-task pairs with attacks ranging from obviously malicious injections to subtle, context-dependent attacks hidden in otherwise legitimate instructions. We evaluate frontier LLMs on SkillInject, measuring both security in terms of harmful instruction avoidance and utility in terms of legitimate instruction compliance. Our results show that today's agents are highly vulnerable with up to 80% attack success rate with frontier models, often executing extremely harmful instructions including data exfiltration, destructive action, and ransomware-like behavior. They furthermore suggest that this problem will not be solved through model scaling or simple input filtering, but that robust agent security will require context-aware authorization frameworks. Our benchmark is available at https://www.skill-inject.com/.
title Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2602.20156