SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yunhao, Ding, Yifan, Tan, Yingshui, Zheng, Boren, Guo, Yanming, Li, Xiaolong, Zhai, Kun, Li, Yishan, Huang, Wenke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917541667930112
author Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Zheng, Boren
Guo, Yanming
Li, Xiaolong
Zhai, Kun
Li, Yishan
Huang, Wenke
author_facet Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Zheng, Boren
Guo, Yanming
Li, Xiaolong
Zhai, Kun
Li, Yishan
Huang, Wenke
contents Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamined security attack surface. We propose SkillTrojan, a backdoor attack that targets skill implementations rather than model parameters or training data. SkillTrojan embeds malicious logic inside otherwise plausible skills and leverages standard skill composition to reconstruct and execute an attacker-specified payload. The attack partitions an encrypted payload across multiple benign-looking skill invocations and activates only under a predefined trigger. SkillTrojan also supports automated synthesis of backdoored skills from arbitrary skill templates, enabling scalable propagation across skill-based agent ecosystems. To enable systematic evaluation, we release a dataset of 3,000+ curated backdoored skills spanning diverse skill patterns and trigger-payload configurations. We instantiate SkillTrojan in a representative code-based agent setting and evaluate both clean-task utility and attack success rate. Our results show that skill-level backdoors can be highly effective with minimal degradation of benign behavior, exposing a critical blind spot in current skill-based agent architectures and motivating defenses that explicitly reason about skill composition and execution. Concretely, on EHR SQL, SkillTrojan attains up to 97.2% ASR while maintaining 89.3% clean ACC on GPT-5.2-1211-Global.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Zheng, Boren
Guo, Yanming
Li, Xiaolong
Zhai, Kun
Li, Yishan
Huang, Wenke
Cryptography and Security
Artificial Intelligence
Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamined security attack surface. We propose SkillTrojan, a backdoor attack that targets skill implementations rather than model parameters or training data. SkillTrojan embeds malicious logic inside otherwise plausible skills and leverages standard skill composition to reconstruct and execute an attacker-specified payload. The attack partitions an encrypted payload across multiple benign-looking skill invocations and activates only under a predefined trigger. SkillTrojan also supports automated synthesis of backdoored skills from arbitrary skill templates, enabling scalable propagation across skill-based agent ecosystems. To enable systematic evaluation, we release a dataset of 3,000+ curated backdoored skills spanning diverse skill patterns and trigger-payload configurations. We instantiate SkillTrojan in a representative code-based agent setting and evaluate both clean-task utility and attack success rate. Our results show that skill-level backdoors can be highly effective with minimal degradation of benign behavior, exposing a critical blind spot in current skill-based agent architectures and motivating defenses that explicitly reason about skill composition and execution. Concretely, on EHR SQL, SkillTrojan attains up to 97.2% ASR while maintaining 89.3% clean ACC on GPT-5.2-1211-Global.
title SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2604.06811