Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
Fuente:
arXiv
Salvato in:
| Autore principale: | Noever, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Models And A Second Opinion Use Case: The Pocket Professional
di: Noever, David
Pubblicazione: (2024)
di: Noever, David
Pubblicazione: (2024)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
di: Noever, David, et al.
Pubblicazione: (2024)
di: Noever, David, et al.
Pubblicazione: (2024)
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
di: Carauleanu, Marc, et al.
Pubblicazione: (2024)
di: Carauleanu, Marc, et al.
Pubblicazione: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
di: Shi, Yunfan
Pubblicazione: (2024)
di: Shi, Yunfan
Pubblicazione: (2024)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
di: Chu, Junjie, et al.
Pubblicazione: (2026)
di: Chu, Junjie, et al.
Pubblicazione: (2026)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
di: Ristea, Dan, et al.
Pubblicazione: (2024)
di: Ristea, Dan, et al.
Pubblicazione: (2024)
SkillTester: Benchmarking Utility and Security of Agent Skills
di: Wang, Leye, et al.
Pubblicazione: (2026)
di: Wang, Leye, et al.
Pubblicazione: (2026)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
di: Jia, Xiaojun, et al.
Pubblicazione: (2026)
di: Jia, Xiaojun, et al.
Pubblicazione: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
di: Holzbauer, Florian, et al.
Pubblicazione: (2026)
di: Holzbauer, Florian, et al.
Pubblicazione: (2026)
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
di: Hou, Yinghan, et al.
Pubblicazione: (2026)
di: Hou, Yinghan, et al.
Pubblicazione: (2026)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
di: Zhuang, Haomin, et al.
Pubblicazione: (2026)
di: Zhuang, Haomin, et al.
Pubblicazione: (2026)
Detecting Adversarial Fine-tuning with Auditing Agents
di: Egler, Sarah, et al.
Pubblicazione: (2025)
di: Egler, Sarah, et al.
Pubblicazione: (2025)
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2026)
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2026)
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
di: Lv, Lijia, et al.
Pubblicazione: (2026)
di: Lv, Lijia, et al.
Pubblicazione: (2026)
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
di: Li, Ying, et al.
Pubblicazione: (2026)
di: Li, Ying, et al.
Pubblicazione: (2026)
Living Off the LLM: How LLMs Will Change Adversary Tactics
di: Oesch, Sean, et al.
Pubblicazione: (2025)
di: Oesch, Sean, et al.
Pubblicazione: (2025)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
di: Li, Zhiyuan, et al.
Pubblicazione: (2026)
di: Li, Zhiyuan, et al.
Pubblicazione: (2026)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
di: Zhou, Zhaojiacheng
Pubblicazione: (2026)
di: Zhou, Zhaojiacheng
Pubblicazione: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
di: Domico, Kyle, et al.
Pubblicazione: (2025)
di: Domico, Kyle, et al.
Pubblicazione: (2025)
SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents
di: Ouyang, Yipeng, et al.
Pubblicazione: (2026)
di: Ouyang, Yipeng, et al.
Pubblicazione: (2026)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
di: Chen, Zhihao, et al.
Pubblicazione: (2026)
di: Chen, Zhihao, et al.
Pubblicazione: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
di: Panebianco, Francesco, et al.
Pubblicazione: (2025)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
di: Liu, Songyang, et al.
Pubblicazione: (2026)
di: Liu, Songyang, et al.
Pubblicazione: (2026)
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
DECEPTICON: How Dark Patterns Manipulate Web Agents
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
di: Shao, Shuo, et al.
Pubblicazione: (2024)
di: Shao, Shuo, et al.
Pubblicazione: (2024)
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
di: Kang, Daewon, et al.
Pubblicazione: (2025)
di: Kang, Daewon, et al.
Pubblicazione: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
di: Wang, Su, et al.
Pubblicazione: (2026)
di: Wang, Su, et al.
Pubblicazione: (2026)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
di: Afane, Khalifa, et al.
Pubblicazione: (2024)
di: Afane, Khalifa, et al.
Pubblicazione: (2024)
Behavioral Integrity Verification for AI Agent Skills
di: Wu, Yuhao, et al.
Pubblicazione: (2026)
di: Wu, Yuhao, et al.
Pubblicazione: (2026)
AccLock: Unlocking Identity with Heartbeat Using In-Ear Accelerometers
di: Wang, Lei, et al.
Pubblicazione: (2026)
di: Wang, Lei, et al.
Pubblicazione: (2026)
From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection
di: Wang, Haowei, et al.
Pubblicazione: (2024)
di: Wang, Haowei, et al.
Pubblicazione: (2024)
E-PhishGen: Unlocking Novel Research in Phishing Email Detection
di: Pajola, Luca, et al.
Pubblicazione: (2025)
di: Pajola, Luca, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Language Models And A Second Opinion Use Case: The Pocket Professional
di: Noever, David
Pubblicazione: (2024) -
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
di: Noever, David, et al.
Pubblicazione: (2024) -
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
di: Carauleanu, Marc, et al.
Pubblicazione: (2024) -
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026) -
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
di: Shi, Yunfan
Pubblicazione: (2024)