Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmotz, David, Beurer-Kellner, Luca, Abdelnabi, Sahar, Andriushchenko, Maksym |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
von: Schmotz, David, et al.
Veröffentlicht: (2025)
von: Schmotz, David, et al.
Veröffentlicht: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
von: Das, Debeshee, et al.
Veröffentlicht: (2025)
von: Das, Debeshee, et al.
Veröffentlicht: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2026)
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2026)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
von: Hsu, Chia-Yi, et al.
Veröffentlicht: (2026)
von: Hsu, Chia-Yi, et al.
Veröffentlicht: (2026)
Design Patterns for Securing LLM Agents against Prompt Injections
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2025)
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2025)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
von: Jin, Chang, et al.
Veröffentlicht: (2026)
von: Jin, Chang, et al.
Veröffentlicht: (2026)
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
von: Das, Debeshee, et al.
Veröffentlicht: (2026)
von: Das, Debeshee, et al.
Veröffentlicht: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
von: Bertran, Martin, et al.
Veröffentlicht: (2024)
von: Bertran, Martin, et al.
Veröffentlicht: (2024)
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
von: Kalavasis, Alkis, et al.
Veröffentlicht: (2024)
von: Kalavasis, Alkis, et al.
Veröffentlicht: (2024)
Unveiling ECC Vulnerabilities: LSTM Networks for Operation Recognition in Side-Channel Attacks
von: Battistello, Alberto, et al.
Veröffentlicht: (2025)
von: Battistello, Alberto, et al.
Veröffentlicht: (2025)
On the Abuse and Detection of Polyglot Files
von: Koch, Luke, et al.
Veröffentlicht: (2024)
von: Koch, Luke, et al.
Veröffentlicht: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
Revisiting Label Inference Attacks in Vertical Federated Learning: Why They Are Vulnerable and How to Defend
von: Liu, Yige, et al.
Veröffentlicht: (2026)
von: Liu, Yige, et al.
Veröffentlicht: (2026)
SoK: Reducing the Vulnerability of Fine-tuned Language Models to Membership Inference Attacks
von: Amit, Guy, et al.
Veröffentlicht: (2024)
von: Amit, Guy, et al.
Veröffentlicht: (2024)
Sponge Attacks on Sensing AI: Energy-Latency Vulnerabilities and Defense via Model Pruning
von: Hasan, Syed Mhamudul, et al.
Veröffentlicht: (2025)
von: Hasan, Syed Mhamudul, et al.
Veröffentlicht: (2025)
On the Role of Similarity in Detecting Masquerading Files
von: Oliver, Jonathan, et al.
Veröffentlicht: (2024)
von: Oliver, Jonathan, et al.
Veröffentlicht: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
Moshi Moshi? A Model Selection Hijacking Adversarial Attack
von: Petrucci, Riccardo, et al.
Veröffentlicht: (2025)
von: Petrucci, Riccardo, et al.
Veröffentlicht: (2025)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Adversarial Vulnerabilities in Neural Operator Digital Twins: Gradient-Free Attacks on Nuclear Thermal-Hydraulic Surrogates
von: Roy, Samrendra, et al.
Veröffentlicht: (2026)
von: Roy, Samrendra, et al.
Veröffentlicht: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
Vulnerabilities in Machine Learning-Based Voice Disorder Detection Systems
von: Perelli, Gianpaolo, et al.
Veröffentlicht: (2024)
von: Perelli, Gianpaolo, et al.
Veröffentlicht: (2024)
Evasion of IoT Malware Detection via Dummy Code Injection
von: Zargarzadeh, Sahar, et al.
Veröffentlicht: (2026)
von: Zargarzadeh, Sahar, et al.
Veröffentlicht: (2026)
Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents
von: Lukáš, Ondřej, et al.
Veröffentlicht: (2026)
von: Lukáš, Ondřej, et al.
Veröffentlicht: (2026)
MRMMIA: Membership Inference Attacks on Memory in Chat Agents
von: Chen, Kai, et al.
Veröffentlicht: (2026)
von: Chen, Kai, et al.
Veröffentlicht: (2026)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
von: Schmotz, David, et al.
Veröffentlicht: (2025) -
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024) -
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
von: Das, Debeshee, et al.
Veröffentlicht: (2025) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024) -
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2026)