HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yukun, Zhang, Yage, Backes, Michael, Shen, Xinyue, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026)
by: Zhang, Yage, et al.
Published: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
by: Jia, Xiaojun, et al.
Published: (2026)
by: Jia, Xiaojun, et al.
Published: (2026)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
SkillTester: Benchmarking Utility and Security of Agent Skills
by: Wang, Leye, et al.
Published: (2026)
by: Wang, Leye, et al.
Published: (2026)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
by: Hou, Yinghan, et al.
Published: (2026)
by: Hou, Yinghan, et al.
Published: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
by: Yang, Fan
Published: (2025)
by: Yang, Fan
Published: (2025)
Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm
by: Marino, Bill, et al.
Published: (2025)
by: Marino, Bill, et al.
Published: (2025)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
by: Chen, Zhihao, et al.
Published: (2026)
by: Chen, Zhihao, et al.
Published: (2026)
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
by: Rahman, Md Hafizur, et al.
Published: (2026)
by: Rahman, Md Hafizur, et al.
Published: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents
by: Ouyang, Yipeng, et al.
Published: (2026)
by: Ouyang, Yipeng, et al.
Published: (2026)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
by: Noever, David
Published: (2025)
by: Noever, David
Published: (2025)
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
by: Das, Debeshee, et al.
Published: (2026)
by: Das, Debeshee, et al.
Published: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
by: Beurer-Kellner, Luca, et al.
Published: (2026)
by: Beurer-Kellner, Luca, et al.
Published: (2026)
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
by: Lv, Lijia, et al.
Published: (2026)
by: Lv, Lijia, et al.
Published: (2026)
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
by: Li, Ying, et al.
Published: (2026)
by: Li, Ying, et al.
Published: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts
by: Siddiky, Md Nurul Absar
Published: (2026)
by: Siddiky, Md Nurul Absar
Published: (2026)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
by: Zhou, Zhaojiacheng
Published: (2026)
by: Zhou, Zhaojiacheng
Published: (2026)
Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
by: Holzbauer, Florian, et al.
Published: (2026)
by: Holzbauer, Florian, et al.
Published: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
by: Xiao, Wenjie, et al.
Published: (2026)
by: Xiao, Wenjie, et al.
Published: (2026)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
Similar Items
-
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
by: Jiang, Yukun, et al.
Published: (2026) -
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026) -
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026) -
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026) -
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
by: Jia, Xiaojun, et al.
Published: (2026)